Auto-publish to the bank
Expanded into a full bank entry and queued to send with no human touch, but only if it clears the high bar on its path (REAL), carries no critical flag, and passed dedup.
Read below for the steps and assumptions in the weekly research/headlines pipeline:
The Research Integrity Gate is not original. It adapts three open peer-review skills built for the Claude ecosystem, each credited and linked so a skeptical reader can check the source.
What is ours is the adaptation: scoping the gate to Tier 1 research only, weighting it toward decision-relevance for an education leader, and wiring its output to a public scored page with CUT verdicts kept internal.
38 sources, 6 tiers, chosen once for trust and not for breadth.
The source list is the first act of judgment, so a clean intake matters more than a large one. The pipeline scrapes it on a 7-day window, weekly, and collects candidates from the trailing week. Each source starts in a tier by how much trust it carries before any scoring. Research and adopted policy start high. Lab and vendor announcements are treated as claims to be tested, never as signal on their own. A few high-trust sources bot-block raw requests, so those are fetched with the browser tools rather than a plain HTTP call.
An LLM compares by subject and claim, not by string match.
Wording shifts as different sources cover the same story, so a string match misses the overlap. Instead, an LLM runs a semantic comparison of each surviving candidate against three things: every issue already sent, everything in the bank waiting to send, and the scored research list. Near-duplicates are dropped or merged into the existing entry, and the call is logged so the merge can be reviewed later.
Does this change what a leader should do, believe, or stop believing about AI in their schools?
Path A · The Research Integrity Gate. Tier 1 papers and studies only. Each is scored 0 to 100 across nine weighted dimensions. Expand any dimension to read the exact criterion the pipeline scores against, taken verbatim from the ranking prompt.
Is this an RCT or a strong quasi-experimental design (diff-in-diff, RD, IV, matched comparison) with confounds controlled, or is it merely correlational? Reward clean causal identification; demote "X is associated with Y" framing dressed up as causal.
Would a principal, district leader, or department head actually change a decision because of this? Inert-but-interesting scores low here too.
Adequate N, representative sample, and adequately powered to detect the claimed effect. Underpowered or convenience samples score low.
Is the effect big enough to matter in a real school, not just statistically detectable? A significant-but-trivial effect scores low.
Watch for p-hacking, uncorrected multiple comparisons, and garden-of-forking-paths signals (many outcomes, flexible specifications, results that hinge on one cut of the data). Pre-specified, robust analyses score high.
Preprint scores lower than peer-reviewed; WWC-vetted scores highest. Preregistration adds points; an undisclosed analysis path loses them.
Does it replicate or converge with the existing body of evidence, or is it a lone surprising result? Reward convergence; treat outliers with caution.
Who funded and authored it? Vendor-funded or self-funded studies of the funder's own product are demoted hard. Independent funding scores high.
Does the finding transfer to a working leader's context, to real classrooms, real constraints, and real populations, or only to a lab or a single atypical site?
Then a devil's advocate tries to refute the study and names the single strongest counter-argument. Any flaw it rates critical, a fatal confound, an undisclosed vendor conflict, a fatally unrepresentative sample, or clear p-hacking, caps the total and vetoes a passing verdict outright. The same scores feed the public Research page.
75 or higher, no critical flag. Holds up.
50 to 74, or a major flaw. Direction is real.
Below 50, or any critical flag. Held back.
Path B · The four-axis score. Journalism, policy, lab announcements, and practitioner items. Each axis runs 0 to 10, then sums. Marketing and items whose only signal is that they are trending are demoted, and vendor business news (funding rounds, valuations, acquisitions, sign-up milestones) is cut outright, because adoption is not evidence.
After scoring, the devil's-advocate pass, and dedup, every surviving item is routed automatically. There is no required per-item human reviewer. The split is rule-based.
Expanded into a full bank entry and queued to send with no human touch, but only if it clears the high bar on its path (REAL), carries no critical flag, and passed dedup.
Everything else: borderline scores, anything carrying a critical flag, and anything touching lesson planning, the conflict zone, which always lands here regardless of score. These wait for a batch audit before they can send.
A separate weekly job pulls the next unused bank entry (oldest first, skipping anything past its refresh date), renders it into the brand template, and sends. No human in this loop. The bank is the buffer that lets a weekly send and a high, fixed bar coexist: when a week is quiet and nothing clears the bar, the bank carries the send. A quiet week that produces nothing is correct behavior, not failure. The bank buffers the timing so the cadence never forces the bar down.