Agent workflows
Before an Agent Charts Your CSV, Require a Reconciliation Gate
A practical CSV-to-report workflow that checks identifiers, schema, missing values and row accounting before an agent turns data into confident-looking charts.
A polished chart can hide an ingestion mistake
A reporting agent receives a monthly CSV and returns a tidy dashboard. The labels are readable and the chart totals look plausible. That does not establish whether customer identifiers retained their leading zeros, missing amounts became zeros, or malformed records disappeared before aggregation.
For a recurring report, the first artifact should be an ingestion and reconciliation note. The chart is downstream of that note. This guide proposes a release gate for teams using agent skills to turn local tabular exports into stakeholder reports.
Methodology: inspect the transformation boundary
This is a documentation-based workflow checked on September 10, 2026, using three DuckDB documentation pages covering CSV detection, import options and aggregates. No private dataset was inspected, and we did not run the proposed fixture or measure error rates.
We chose these boundaries because they separate three questions: what the file contains, how those values are interpreted, and what aggregation means. DuckDB is a concrete reference implementation, not a claim that every reporting skill uses it or that it is the only suitable engine.
Use detection as a hypothesis
The CSV auto-detection documentation explains that DuckDB infers dialect, column types and headers. Detection uses sampling; the documented sniffer can expose the inferred settings separately. Treat those settings as a proposal that needs review against the source system's meaning.
For a customer identifier, a numeric-looking string is not necessarily a quantity. For a date, plausible parsing does not settle whether the producer meant month-first or day-first. Ask for the export contract or a confirmed example instead of asking the agent to guess confidently.
Our recommended first deliverable lists the producer, reporting period, source filename, file identity, detected headers, proposed types and unresolved ambiguities. A full-file scan can reveal values missed by a sample, but it cannot infer a business definition that was never supplied.
Preserve raw values before applying meaning
The CSV import reference documents explicit column types and an option to read columns as text. Those are useful mechanisms when a workflow needs to inspect values before committing to numeric or date interpretations.
Our proposed design keeps the original export unchanged and builds a separate normalized dataset. Preserve an original identifier alongside any derived key. Retain the original amount text when applying a documented numeric conversion. Specify the currency, decimal convention and reporting timezone where relevant; those meanings should come from the data owner.
Reading everything as text is not a finished analysis. It moves interpretation into a visible transformation step. Record conversion failures there. Do not quietly remove difficult rows merely to make a chart render.
Reconcile rows before totals
Define a record as a parsed CSV record, not a physical line: quoted text can contain line breaks. For a file accepted by the parser, our proposed accounting rule is that every input record must have a documented destination, such as accepted, rejected with a reason, or intentionally excluded by an approved reporting rule.
Require mutually exclusive dispositions so the counts can reconcile. If parsing itself fails before reliable record accounting is possible, block the normal report and deliver a diagnostic instead. Do not present an incomplete count as the full export.
For duplicates, ask what identifies a transaction. Two rows sharing a customer number may represent different purchases. A blanket deduplication step can erase valid activity. Preserve suspected duplicates for review until the intended key and policy are known.
Keep missing, zero and empty separate
The aggregate reference documents how general aggregates handle nulls, including that a sum over an empty group is null rather than zero. That behavior matters when a dashboard is expected to distinguish no observed value from an observed value of zero.
Our recommendation is to expose coverage beside each important total. Show the reporting period, included-record count and missing-value count, with an explicit exclusion rule. A numeric result without its eligible population is easy to misread.
Do not automatically replace missing values with zero for presentation convenience. If the business owner approves that rule, record it in the report notes and preserve the original missingness in the normalized data. The decision is semantic, not just a formatting preference.
A small acceptance fixture
Before adopting a reporting skill, propose a non-sensitive fixture containing an identifier with a leading zero, a missing amount, a genuine zero, an ambiguous date and a repeated customer with different transactions. Add one malformed record in a separate parser-failure fixture.
Write expected dispositions with the data owner before running the agent. The acceptance question is not whether the output looks persuasive. It is whether the workflow preserves identifiers, identifies uncertainty, reconciles records and refuses to claim completeness after a parse failure.
Then inspect the report's narrative. A chart showing a change does not establish its cause. Require the agent to separate calculated observations from possible explanations, and to label any causal claim that needs additional evidence.
Limitations and the final handoff
This gate does not catch every upstream data defect, incorrect business definition or privacy problem. A perfectly reconciled export may still omit transactions before export. Local processing also does not guarantee that an agent or connected service never transmits data; inspect the actual workflow permissions.
Hand off the source identity, transformation rules, accepted and rejected record counts, coverage notes and chart inputs together. Keep sensitive row samples out of public debugging logs.
Our earlier CSV report skill introduction discusses presentation-oriented output. This article supplies the validation boundary that should precede such output. Browse the skill directory for candidates, but choose a reporting workflow by its data contract before choosing its visual theme.