Notes from building Good News Agent, a LangGraph workflow that researches, qualifies and writes a personalised newsletter—and from the failures that only appeared once it ran on real sources every day.

Explore this article visually3 figures

01

A duplicate that changes its headline

02

The demo worked on a Tuesday

Good News Agent is a LangGraph graph with five nodes: plan the topics, collect candidate articles, qualify the sources, verify the claims that will be quoted, write the issue. On the first evening the demo produced a convincing newsletter from a handful of feeds. I was pleased with it for about a week.

Then it ran daily. The qualification node scored a paywalled stub as a strong source because the two visible sentences were well written. The same wire story appeared twice under different headlines and both copies were selected. A link resolved during collection and was dead by the time the writer node cited it. None of these are model failures in the usual sense—the model did exactly what a fluent reader would do with the information it had. They are workflow failures: the wrong information was crossing the boundary between nodes.

That reframing changed how I spent the rest of the project. Instead of tuning prompts, I spent the time on what each handoff was allowed to contain.

03

What a handoff should carry

Figure 01Architecture

An evidence record, opened up

Select an element to explore its role.
Read every explanation
Identity
Keep the original address and retrieval time. A stable identifier lets later steps refer to the same source even when a headline changes.
Evidence
Store the retrieved text and the passage behind the claim. A hash detects identical normalised text; it neither archives the page nor detects every near-duplicate.
Decision
Record acceptance or rejection with a short reason and its author. This makes a questionable selection inspectable instead of hiding it inside a final paragraph.
01 / 03
Identity

Keep the original address and retrieval time. A stable identifier lets later steps refer to the same source even when a headline changes.

Proposed handoff schema. Identity, evidence and review state solve different problems; a URL alone cannot replace them.

The state object between nodes started as a list of URLs and free-text summaries. It ended as a typed record where every field exists because a specific failure needed it:

pythonThe candidate record that travels between graph nodes.
class Candidate(TypedDict):
    url: str
    fetched_at: str          # the writer cites this snapshot, not "the page"
    content_hash: str        # dedup key — the title is useless for this
    excerpt: str             # the exact text a claim rests on (≤ 600 chars)
    source_score: float
    reasons: list[str]       # why the qualifier scored it; shown to the reviewer
    status: Literal["candidate", "qualified", "verified", "rejected"]

Three of those fields did most of the work. content_hash is computed on the normalised body text, so two headlines for one wire story collapse into one candidate. excerpt forces the verification node to point at a sentence rather than assert that a page "supports" a claim. reasons is a list of short strings the qualifier must fill in, which turned out to be the cheapest way to see when its scoring was wrong: when a paywall stub scored highly, the reasons said well-structured argument about two sentences, and that was enough to add a length floor.

The rule I ended up with is simple: a node that cannot fill a required field returns the record with status = rejected and a reason. It never writes a paragraph explaining that it could not find something. Prose at a boundary is where debugging goes to die.

04

Where the person actually belongs

The first version put a person at the end, reading the finished newsletter. That is the worst place for review: the draft is polished, the sources are three steps back, and the only available action is to reject the whole thing.

The version that worked put the review between qualification and writing. The reviewer sees a packet of five to eight candidates, each with its excerpt, its source score and its reasons. Accepting or dropping a candidate takes a few seconds because the evidence is already beside it. Every decision is written back into the run's trace with the reviewer's label, which is how the regression set grew without anyone sitting down to write test cases.

Diagram 02

Who holds the evidence at each step

PersonPlannerTools / sourcesEvidence log01020304050607brief + stop conditiontyped queriessources + idsdraft linked to sourcesreview packetaccept / editregression caseThe evidence log receives everything before the person does — review reads sources, not a paragraph.
A swimlane view of one issue of the newsletter. The evidence log is written to before the person ever sees a draft, so review happens on sources and reasons rather than on a fluent paragraph.

One consequence I did not anticipate: once the review moved earlier, the writer node stopped needing a large model. It receives verified excerpts and a structure; it is doing composition, not judgment. The expensive calls are in qualification and verification, where being wrong is costly.

05

Evaluate the trajectory, not the newsletter

Figure 03Feedback loop

Turn one failure into a regression case

Select an element to explore its role.
Read every explanation
Capture
Save the inputs, source snapshots, model configuration and tool results needed to understand the failure.
Label
A reviewer identifies the missed duplicate or unsupported claim and writes the expected behaviour. Fluent output is not the label.
Compare
Replay matched cases and compare source selection as well as output. Investigate changed tool arguments before reusing recorded results.
Keep
Keep the case in the regression suite. A later prompt change must not quietly reintroduce the original failure.
The result informs the next iteration
01 / 04
Capture

Save the inputs, source snapshots, model configuration and tool results needed to understand the failure.

A proposed learning loop for the workflow, not automatic retraining of the model.

A rubric score on the final text told me almost nothing. Two issues could read equally well while one had cited a dead link and dropped the best source. The measurements that actually moved decisions were all about the path:

MetricWhat it caught
Duplicates per issue (by content_hash)The wire-story problem; measure exact duplicates separately from near-duplicates
Excerpt coverage — % of cited claims with a verified excerptThe writer inventing a bridging sentence between two sources
Reviewer drop rate on qualified candidatesThe paywall-stub scoring; investigate changes in source mix and reviewer agreement before adjusting thresholds
Dead links at write timeKeep the evidence snapshot and check the public link separately
Cost per issue by nodeCompare spend with errors prevented; no node has a universal target share

This is also why I later built flightrec, a small recorder and replayer for agent runs. Recording every model call, tool call and state transition as events means a run can be replayed against a new prompt or model and the two trajectories diffed. "Did the new prompt change which sources got qualified?" becomes a question with an answer instead of an impression.

06

What I would do from the start next time

  • Write the state type before the first node. If a field cannot be named, the handoff is not understood yet.
  • Put the review packet where the evidence is freshest, not where the output is prettiest.
  • Log reasons as short strings, not explanations. They are for grep, not for reading.
  • Cite snapshots. The web changes between collection and writing, even within one run.
  • Measure the path from day one; the final-text score is the last metric to add, not the first.

07

A hash is not an evidence archive

A hash identifies an exact normalised body; it does not detect edited or syndicated near-duplicates. Keep a stable source identifier and compare similar passages separately. A timestamp also does not preserve a page: store the retrieved text and its provenance, subject to access and retention rules. Recheck public citation links before publication, and run a final claim-to-excerpt review after writing. Early source approval cannot catch a new claim introduced by the writer.

Discuss this article

More articles

09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read02Designing AI that works without a networkEdge AI · 8 min read