Agent workflows
Your GitHub Agent Finished Paging. Did It Find Everything?
Separate a completed page walk from complete discovery. Give GitHub collection agents a coverage receipt, resumable checkpoints and an honest stopping rule.
The last page is not a census
A GitHub discovery agent returns a spreadsheet of repositories and announces that collection is complete. Its loop reached the last page without throwing an exception. That sounds reassuring, but the output may still be a bounded search result rather than an inventory of everything relevant.
For teams building a skill directory, dependency inventory or repository research brief, the useful deliverable is not just a list. It is a list plus an explanation of its coverage. A stopped loop, an exhausted query and a complete research population are different claims.
Methodology: three boundaries, not a speed contest
This documentation-based guide was checked on September 22, 2026. We selected GitHub's pagination, search and REST best-practice documentation because they describe transport, discovery limits and request behavior respectively. We did not run a large collection job or measure throughput.
The workflow below is our proposed engineering handoff. It is intended for read-only collection agents, not for authorizing repository changes. Our earlier browser-versus-crawler guide helps choose an extraction mechanism; this guide addresses what a completed API collection may legitimately claim.
Follow the server's continuation, not a guessed page count
GitHub's pagination documentation describes following the response's Link header. Available relations vary; a last-page link is not always supplied. Endpoints also differ in their pagination parameters.
Our recommendation is to preserve the next request supplied by the service rather than reconstructing it from an assumed numbering scheme. Record the endpoint, query, sort settings, API version and access scope alongside that continuation. Do not record authentication secrets.
Choose page-at-a-time persistence when the result could exceed the agent's context or process memory. The document describes an iterator as well as an aggregate pagination helper. A helper that collects pages is convenient, but it should not become the only record of where the job stopped.
Separate search exhaustion from discovery coverage
GitHub's search reference specifies a maximum of 1,000 returned results per search. It also documents the incomplete_results signal for time-limited queries and differences caused by resource access.
Consequently, reaching the end of a search does not prove that every matching repository was returned. Nor does an apparently complete result establish access to private resources that the caller cannot see.
For our proposed coverage receipt, keep these fields separate: requested population, authenticated visibility scope, query partitions, unique repository count, returned completeness signals and unresolved gaps. Use a label such as bounded discovery when the collection method cannot justify a census. Do not rename that state complete merely because enough interesting candidates were found.
If you partition a search into narrower queries, document the boundaries and deduplicate the combined results. Narrower queries are a coverage technique, not permission to evade quotas. They also do not automatically establish a stable snapshot while repositories are changing.
Make restart behavior explicit
Consider a proposed collector that saves a page of repository records and then crashes before saving its continuation. On restart, that page may be fetched again. The persistence layer should tolerate the replay rather than appending duplicate public entries.
Use the repository's stable identifier as the collection key, with observed names and URLs as attributes. Keep a separate skill identity when one repository contains several skills; repository deduplication alone cannot distinguish those packages.
Commit records and the next checkpoint together where practical. Otherwise, design for replay and record which operation completed. A checkpoint should identify the run and partition, not just a bare page number detached from its query. Never mark a failed request as an empty final page.
Give the agent a bounded stopping policy
GitHub's REST best practices recommend serial requests to reduce secondary-limit pressure. They also explain respecting Retry-After, waiting for reset when the remaining primary allowance is zero, and backing off on repeated secondary-limit failures.
Our additional recommendation is a run budget covering elapsed time, pages and retries. When that budget is exhausted, persist the checkpoint and return partial with a reason. Do not open more workers or rotate credentials to force completion.
Separate the collection scheduler from the publication scheduler. Finding a repository is not evidence that its skill can be installed safely, and a resumed discovery run should not silently republish an existing listing.
A failure rehearsal before daily scheduling
Use a local mock service and non-sensitive fixtures for four proposed acceptance cases:
- A response supplies a next link, followed by a normal final page.
- Search reports a larger population than the query can return.
- A rate-limit response interrupts the walk without returning records.
- Persistence succeeds, but the worker stops before advancing its checkpoint.
For each case, write the expected run status, record count and restart behavior before executing the collector. Check that the narrative matches the state: an interrupted job must not report zero new repositories as a complete scan. These are suggested tests, not results from an OpenAgentSkill production experiment.
Limitations and the useful handoff
A coverage receipt cannot recover records that the upstream search never exposed. Deduplication cannot restore items skipped by a changing ordering, and a successful restart is not proof of snapshot consistency. State those limitations explicitly when delivering research.
When choosing a workflow from the skill directory, ask it to return its stopping reason and checkpoint contract before asking how many repositories it can scan per day. The stronger answer is an auditable boundary around the result, not the largest unsupported number.