Back to Blog

Agent workflows

Give Your Design Agent the Image's Job Before Asking for Alt Text

The same image can need different text in a gallery, a button or a report. Review context, accessible names and chart explanations before bulk-generating descriptions.

OpenAgentSkillPublished:

One asset, several different jobs

A design agent receives a folder of images and writes a description for each file. The descriptions sound fluent. Then the same illustration appears in a decorative banner, a gallery card and an image-only link. Reusing one sentence everywhere can create repetition in one place and conceal the intended action in another.

The missing input is the image's job in the page. For teams shipping agent-generated interfaces, alternative text should be reviewed per placement, not approved once at the asset-library level.

Methodology: review meaning in context

This guide uses three W3C Web Accessibility Initiative tutorials, checked on September 22, 2026: the alt decision tree, functional images and complex images. They were selected to cover ordinary content, interactive use and information that cannot fit into a short description.

The placement record and exercises below are our editorial recommendations. We did not conduct a screen-reader study, certify a component or measure task completion. The intended audience is a designer or frontend maintainer reviewing an agent's implementation, rather than an image-captioning model evaluation team.

Start with a placement record

The WAI alt decision tree makes context central: an image may communicate information, duplicate nearby text, perform a function or be purely decorative. Its guidance includes empty alt text for appropriate decorative or redundant cases.

Our proposed handoff records the page route, component, image identifier, surrounding copy, interaction and intended contribution. Add the proposed text alternative and a short reason. Keep an unresolved state when the agent cannot determine the role.

For example, a paper-texture image behind a heading may add atmosphere without contributing information. A close-up used to demonstrate print texture in a design case study has a different purpose. The source file alone cannot decide which treatment is appropriate.

This is also why a bulk rule requiring nonempty descriptions on every image is a poor acceptance criterion. Check the correctness of the decision, not simply whether a string exists.

Name the action when the image is the control

WAI's functional-image tutorial distinguishes the purpose of a linked or button image from its appearance. It also shows how nearby link text can make an image description redundant.

Consider a proposed gallery card with a preview and a separate play button. Describing the button as a white triangle says little about its action. Naming it Play preview identifies the operation; the artwork description can remain associated with the artwork.

Now consider a card where the image and visible title belong to the same link. Review the resulting accessible name as a whole. Do not assume that adding the title again to the image improves the experience. Conversely, removing a duplicate description must not remove the only meaningful name from an image-only link.

Ask the agent to show the rendered structure, not just its proposed copy. Whether text belongs to the same control matters. A screenshot cannot establish that relationship.

Treat charts as explanations, not caption contests

The WAI complex-image tutorial describes short identification paired with a longer textual equivalent for substantial visual information. Its examples include descriptions containing values, relationships and trends.

Our recommendation for an agent-generated report is to attach the explanation to the same approved data revision as the chart. If the chart changes, the text needs another review. An elegant description of last month's values is still wrong.

A proposed chart handoff should name the question the chart answers, the population, units, period and important comparisons. Provide a readable data table where that serves the task, and explain relationships that a table alone may not communicate. Do not ask a vision model to guess small labels or invent precise values from an ambiguous screenshot.

When source values are unavailable, say so and request them. A generic phrase such as performance chart cannot replace the evidence that the visualization was meant to communicate.

Review a component in three placements

Create a non-production example using one authorized product image. Place it in a decorative header, an image-only destination link and a detail-page figure that explains a specific feature. Write the expected purpose of each placement before the agent generates text.

Inspect the page with images unavailable, then inspect the accessibility tree and keyboard navigation. Where appropriate, include a screen-reader check with someone familiar with the target interaction. Record which checks were actually completed and which remain unverified.

For a chart, revise one underlying value and confirm that both visual and textual outputs are flagged for review. For a card, change its destination and check the control name. These exercises test whether content remains coupled to purpose; they are not claims that any listed skill already passes.

Keep localization attached to the component

Our suggested content model stores translations with the placement record. Translating an asset description once is insufficient when the component's function or nearby text changes between layouts.

During review, ask whether the localized wording communicates the same action or information without unnecessary duplication. Preserve product names when needed, but do not fill alternative text with search keywords or promotional claims that are absent from the image's role.

The OpenAgentSkill Gallery can supply creative references. A visually compelling example is a starting point for design discussion, not proof that its accessibility behavior suits your own page.

Limitations and the final review

This workflow does not cover every accessibility requirement, every kind of graphic or every assistive-technology combination. It cannot determine whether an ambiguous image is factually accurate, licensed or appropriate for the audience. Those remain separate reviews.

Use the form error-recovery guide for another interaction-specific review, and browse the skill directory for implementation candidates. Ask the selected agent for an image-role inventory alongside its visual design. Good alternative text begins with knowing what the interface needs the image to do.