Methodology

Read below for the steps and assumptions in the weekly research/headlines pipeline:

The system to date
38
Curated sources
6
Source tiers
14/15
Studies cleared the bar
9
Dimensions per study · 0 to 100
1
Issue published
The methods we adapt

The scoring method is adapted from open peer-review work

The Research Integrity Gate is not original. It adapts three open peer-review skills built for the Claude ecosystem, each credited and linked so a skeptical reader can check the source.

What is ours is the adaptation: scoping the gate to Tier 1 research only, weighting it toward decision-relevance for an education leader, and wiring its output to a public scored page with CUT verdicts kept internal.

Does a human review every issue?

A human does NOT review every issue. My plan is to give the digest a once-over before sending, but the goal is that it will eventually be fully automated. If there appears to be faulty judgement, I will look into the prompts that build toward each issue. In summary, the judgment is built into the filter, then audited in batches.

The pipeline

5 steps executed weekly:

Detect › Dedup › Rank + Gate › Publish or Hold › Send

A curated credibility list, not a wide net

38 sources, 6 tiers, chosen once for trust and not for breadth.

The source list is the first act of judgment, so a clean intake matters more than a large one. The pipeline scrapes it on a 7-day window, weekly, and collects candidates from the trailing week. Each source starts in a tier by how much trust it carries before any scoring. Research and adopted policy start high. Lab and vendor announcements are treated as claims to be tested, never as signal on their own. A few high-trust sources bot-block raw requests, so those are fetched with the browser tools rather than a plain HTTP call.

Tier 1 · ResearcharXiv, RAND, AERA, IES / WWC, journals
primary, peer-reviewed
Tier 5 · Policy & govUS OET, state trackers, UK / EU bodies
adopted policy is a fact
Tier 6 · Unions & civil societyunions, privacy / civil-society orgs
shapes contracts, surfaces harms
Tier 2 · JournalismEdWeek, Hechinger, Chalkbeat, The 74
reaching classrooms
Tier 4 · Practitionerssmall, named, non-promotional
track record, reviewed often
Tier 3 · Labs / vendorsOpenAI, Anthropic, Google, Khan
claims to test
Bar width = starting credibility weight before any scoring. A vendor's claim about its own product starts near zero. Full list and per-source status live in config/sources.yaml.
The Margin · A weekly AI-in-education digest for education leaders
My Planning Partner · A studio for tomorrow's lesson
Home · Methodology · Research · My Planning Partner · @myplanningpartner