Back to Blog

Agent workflows

When a Skill Updates, Which Review Evidence Can You Keep?

Bind review evidence to the actual skill version, inspect more than SKILL.md, and separate discovery of an update from authorization to install it.

OpenAgentSkillPublished:

Yesterday's review describes yesterday's package

A skill repository changes after a team has reviewed it. The name, author and star count remain familiar, but the new revision includes an additional helper and a different network destination.

The old review is not worthless. It is historical evidence about the version that was inspected. The mistake is presenting it as if it covered the new contents. For catalog operators and teams maintaining installed skills, the useful question is which evidence remains applicable and which claims need another check.

Methodology and intended audience

This is a documentation-based design proposal reviewed on September 11, 2026. We examined the Agent Skills specification, Git's diff documentation and GitHub's release documentation. We did not execute a third-party skill, inspect a suspicious repository or change any production review status.

The workflow is meant for teams designing update handling. It is not a statement that OpenAgentSkill or every agent host already implements the proposed fields. Its selection criteria are evidence scope, reviewability and the ability to explain what changed.

Identify the package boundary

The Agent Skills specification defines a skill directory containing SKILL.md and potentially scripts, references, assets and other resources. Reviewing only the main Markdown file can therefore miss changes in material that the agent later reads or runs.

Our proposed identity record includes the repository, resolved source revision, skill path and the exact package contents delivered to the host. If an installer copies only a subset of the repository, preserve the identity of that subset rather than assuming that a repository-level label fully describes it.

Also list external resources whose contents are not captured by that identity. A helper that downloads a remote script or model introduces a separate dependency boundary. An unchanged local file can still invoke changed remote behavior.

Compare the version that was reviewed with the candidate

The git-diff documentation describes comparing content and reporting changed file names and statuses. A file-level inventory is a useful first pass, but it is not a behavioral review.

For each candidate update, our recommendation is to compare the previous reviewed contents with the intended replacement. Group changes by what they can affect: instructions, executable helpers, dependency declarations, referenced material, configuration and licensing information.

Do not classify Markdown as harmless just because it is not a script. Instructions can direct an agent toward different actions. Likewise, do not declare a change dangerous merely because it introduces a helper. Describe the changed capability, relevant access and remaining uncertainty so the reviewer can make a scoped decision.

Keep historical facts, refresh current claims

Some evidence should remain unchanged: who reviewed the older revision, when they reviewed it, what they examined and what they observed. Preserve that record even when the new candidate is rejected.

Other claims should be recomputed or checked against the candidate. A previous scan result should not automatically describe modified files. A prior installation outcome should not imply that the new version was installed successfully. A creator's identity can remain the same while the contents they maintain change.

Our proposed public presentation distinguishes the current source version from the latest reviewed version. If they differ, say so directly. A useful status is update discovered, review pending, with the earlier review still available as history. That is more informative than either erasing the history or silently carrying its badge forward.

Treat release names as references, not complete evidence

The GitHub release documentation explains that releases are based on tags marking points in repository history and may include release notes and downloadable files.

For an update workflow, record the source revision actually resolved and the identity of the downloaded artifact when one is used. Release notes help a reviewer navigate the change, but they are not a substitute for inspecting what the installer will deliver. Do not assume that every asset has the same contents or permissions as a source checkout.

If the relationship between an artifact and reviewed source cannot be established, keep that uncertainty visible. A familiar version label alone should not turn an unknown artifact into a reviewed package.

Separate notification, review and installation

An update detector can notify a team that a candidate exists without changing the installed copy. Review can inspect that candidate without authorizing its execution. Installation can be approved for a limited workspace without authorizing production access.

Keep those transitions separate in both the interface and automation. Our proposed update receipt records the old version, candidate version, changed capabilities, reviewer decision, intended target and any installation result actually observed.

Where a user has local modifications, preserve them for comparison and conflict handling. A blanket replacement may solve version drift while destroying useful customizations. The update plan should say how those differences will be retained or resolved before files are changed.

A proposed regression exercise

Use a disposable fixture repository with two non-sensitive versions. In one update, change only explanatory wording. In another, add a helper that declares a new external destination. In a third, leave local files unchanged but make an external dependency unresolved.

Write the expected review behavior first. The workflow should preserve historical evidence, expose changed or unknown boundaries and avoid reporting installation success before an installation has occurred. These are suggested tests, not observations from a system we ran.

Limitations and practical adoption

Diffs cannot prove that code is safe, and version identity does not freeze every external service. Runtime permissions, host behavior, dependency resolution and task context still matter. Review evidence should state its scope instead of promising universal safety.

Our Python dependency handoff guide covers one part of that external dependency problem. When selecting workflows from the OpenAgentSkill directory, ask which exact version an audit or outcome describes.

The long-term objective is not to stop updates. It is to make each adoption decision explainable without pretending that an older review inspected code that did not yet exist.