Back to Blog

Agent workflows

Handle HTTP 304 Without Erasing Your Web Agent's Evidence

Design conditional web collection around retained response bodies, request variants and extractor versions. A 304 response should not become an empty record or a false change alert.

OpenAgentSkillPublished:

The empty response that should not empty the database

Consider a web agent monitoring a public documentation page. Yesterday it retained the page and extracted a list of supported integrations. Today it sends a conditional request, receives HTTP 304 and writes an empty integration list because the response has no body.

The collection job has confused a protocol result with new source content. The correct next step is not to ask a model to reconstruct the missing page. It is to locate the appropriate retained representation and decide whether that evidence can still be reused.

This guide is for teams building scheduled, authorized collection of public pages. It focuses on ordinary full-response requests, not resumable downloads, authenticated account scraping or every feature of a general-purpose HTTP cache.

Methodology and the two caches to design

We reviewed MDN's HTTP documentation and the HTTP caching standard on October 3, 2026. The workflow below is a proposed design; we did not operate a production crawler or measure bandwidth savings.

Separate two layers in the design. The source layer retains an HTTP representation with its validation metadata. The extraction layer turns that representation into fields using a particular parser and schema. Reusing the first does not automatically justify reusing the second.

This is a more specific implementation concern than choosing a browser or crawler. Our extraction contract comparison covers that earlier decision.

Interpret 304 before calling an extractor

MDN's 304 reference explains that the server need not retransmit the representation after a qualifying conditional request. The response does not contain a body. It is therefore not a new empty document.

Our recommended collector handles this status before parsing. It finds the retained body associated with the request and validator, records that validation occurred, and follows the applicable cache rules. Only then does it decide whether an existing extracted result remains suitable.

Keep body retrieval time distinct from revalidation time. A report can accurately say that earlier evidence was revalidated today without claiming the body was downloaded again. This distinction also helps explain why two collections have identical stored source bytes but different observation times.

Bind validators to the representation they describe

An ETag is a server-provided validator, not a value the agent should invent from a page title. MDN's ETag documentation distinguishes weak validators from strong ones: weak equivalence does not promise byte-for-byte identity.

For our proposed collector, store the response body reference and validator together. Preserve the value exactly, including a weak prefix when present. Do not treat a matching ETag as a cryptographic integrity certificate or as evidence that every linked resource is unchanged.

If local cleanup deletes a retained body but leaves its validator, the next 304 cannot restore the missing evidence. Mark this as a cache-record inconsistency. Within the source's access and rate limits, make a bounded request without the conditional header to obtain the body again. Do not manufacture an empty successful extraction or loop indefinitely.

Do not mix request variants

A URL alone may not identify the content the collector intended to retrieve. MDN's Vary documentation explains how named request headers affect representation selection and cache reuse.

For a controlled public documentation monitor, define the intended language and other relevant request headers before scheduling. If the server varies on Accept-Language, do not pair an English validator with a stored body from a different language request.

Use a maintained HTTP caching implementation when possible. The HTTP caching standard defines storage and reuse constraints, variant matching and updating stored responses after validation. A few hand-written ETag branches are not a complete implementation of that standard. Respect no-store and other applicable directives; conditional requests are not permission to retain restricted content.

Our scope recommendation is conservative: keep authenticated sessions out of this shared public-page cache. If private collection is required later, design authorization, retention and partitioning separately rather than adding a cookie to the public workflow.

Separate source changes from parser changes

Suppose the retained source is unchanged, but the team corrects an extractor that previously confused a heading with an integration name. Reusing yesterday's parsed result would preserve the bug even though today's source validation succeeded.

We recommend keying a derived result by the retained artifact identity, extractor version and schema version. A changed extractor can rerun against authorized retained evidence without downloading it again. Its resulting field differences should be labeled as a processing change until source evidence establishes otherwise.

Conversely, a newly downloaded page may differ only in presentation while yielding the same requested fields. Record both observations: the source representation changed, but the accepted domain record did not. Neither observation should be silently substituted for the other in a change alert.

Make outcomes explicit

The following is a proposed application policy, not a measured compatibility result:

ObservationCollector actionReport outcome
Successful response with a usable bodyStore when allowed, extract and validate fieldsNew source observation
304 with matching retained evidenceReuse evidence under cache rulesRevalidated source
304 but retained evidence missingAttempt bounded unconditional recoveryIncomplete until recovered
Source reusable but parser version changedReprocess retained evidenceProcessing revision
Access denied, timeout or server failurePreserve prior evidence without marking it currentCollection failed

Add a small fixture set covering each row before scheduling. Include one language variant and one broken extractor revision. Check that failures never overwrite accepted fields with empty values and that revalidation does not hide an extraction failure.

Track request outcomes, body bytes transferred, extraction executions and accepted records separately. Only measured data can establish actual savings. A lower download volume does not necessarily imply less parsing work or lower model spend.

Limitations and the handoff to an agent skill

Some sites omit validators or implement them incorrectly. A page can load its important data from other endpoints, so validating the initial HTML may not validate the information the user requested. Browser-managed caches and service workers can also affect the observations an integration exposes.

Choose the source artifact that actually supports the extracted fields. Do not label the whole website unchanged because one request returned 304. Continue to respect robots policies, access terms, rate limits and your user's collection scope.

A reusable skill from the OpenAgentSkill directory can document these rules, but the application must enforce storage and state transitions. The desired output is an evidence-backed collection record, not merely a successful request counter.

Handle HTTP 304 Without Erasing Your Web Agent's Evidence | OpenAgentSkill