Every study here is run through one rigorous bar before it earns a verdict. Ordered by a blend of how well it holds up, how influential it is, and how recent. Click any row for the full scoring.
| Year | Verdict | Score | Cited | Study | What it means for leaders | |
|---|---|---|---|---|---|---|
| 2026 | Real | 85/100 | n/a |
Allcott, Baron, Dee, Duckworth, Gentzkow & Jacob · NBER Working Paper 35132
Phones, attention & achievement
|
The strongest causal evidence yet that phone restriction changes behaviour and climate but not test scores. Budget and justify a ban on attention and climate over a multi-year horizon, and plan for a worse first year, not on promised achievement gains. | › |
The readA staggered difference-in-differences study of lockable phone pouches nationwide, combining large-scale surveys, GPS pings, standardized test scores, school administrative records and the largest pouch provider's sales records. Pouch adoption substantially reduces phone use. In year one disciplinary incidents rise and student well-being falls; both reverse in later years. Average effects on test scores are consistently close to zero, with modest positive high-school effects (particularly math) and small negative middle-school effects.
Strongest counter-argumentPouch adoption is not randomly assigned, and staggered difference-in-differences can be biased when treatment effects vary across adoption cohorts. The GPS compliance evidence, however, rules out the strongest rival explanation, that the bans were simply never enforced.
Confidence: high. Scored from the NBER abstract and paper listing; design, data sources and directional findings are stated explicitly by the authors. Citation count not shown because the paper is too recent for a stable count.
| ||||||
| 2024 | Real | 76/100 | 513 |
Fan, Tang, Le, et al. · British Journal of Educational Technology
AI, writing & critical thinking
|
Credible randomized evidence that AI help improves the artifact more than the learner. Design assessments and tooling so AI assists the process rather than substituting for it. | › |
The readIn a randomized four-condition experiment, ChatGPT users improved their essay scores the most but gained no more knowledge transfer than other groups, suggesting AI support can boost the output while offloading the thinking.
Strongest counter-argumentThe 'metacognitive laziness' label over-reads a null. ChatGPT users produced better essays and merely showed no advantage on knowledge transfer in one short lab task, and a missing transfer benefit in an underpowered single session is not evidence of cognitive harm.
Confidence: medium. Scored from the abstract and detailed secondary summaries confirming the four-arm design and N=117; full tables not read directly.
| ||||||
| 2025 | Watch | 51/100 | 408 |
Kosmyna, Hauptmann, et al. (MIT Media Lab) · arXiv preprint
AI, writing & critical thinking
|
Provocative and directionally consistent with offloading concerns, but too small and unreviewed to act on. Treat it as a hypothesis to watch, not proof that AI assistance damages thinking. | › |
The readAcross 54 participants writing essays under EEG, LLM users showed the weakest brain connectivity and the poorest recall of their own essays, which the authors frame as accumulating 'cognitive debt.'
Strongest counter-argumentThe viral 'your brain shuts off with ChatGPT' conclusion is statistically fragile: the headline crossover result rests on 18 participants in a non-randomized fourth session, with dozens of EEG comparisons, no peer review, and a single artificial essay task.
Confidence: low. Scored from the abstract and detailed secondary summaries; full statistical detail not read directly.
| ||||||
| 2025 | Watch | 76/100 | 98 |
Kestin, Miller, Klales, et al. · Scientific Reports
AI tutoring vs active learning
|
A well-engineered AI tutor can match or beat good instruction in a controlled lesson, but the evidence is too narrow to justify replacing classroom teaching. Treat it as promising for piloting supplemental practice, not a deployment mandate. | › |
The readA crossover RCT of 194 Harvard physics students found a custom GPT-4 tutor with engineered pedagogical prompts produced post-test gains of about 0.63 SD, up to 0.73 to 1.3 SD by quantile, highly significant, in less instructional time than in-class active learning.
Strongest counter-argumentThe result comes from 194 elite Harvard physics students on just two topics over roughly 50 minutes, so novelty effects, ceiling effects, and the narrow setting make it unsafe to assume the same gains in a typical K-12 or community-college classroom.
Confidence: high. Read the full published paper including methods, effect sizes, and the no-competing-interests statement.
| ||||||
| 2025 | Watch | 72/100 | 128 |
Letourneau, Deslandes Martineau, et al. · npj Science of Learning
Evidence base (review)
|
The best single map of where ITS evidence is strong versus thin in K-12. The strategic nuance for a buyer: the gain comes from structured tutoring, not necessarily the AI layer, since ITS rarely beat simpler tutoring software. | › |
The readA peer-reviewed synthesis of 28 studies (4,597 K-12 students) finding intelligent tutoring systems generally beat human-teacher-only conditions (Hedges g roughly 0.68 to 1.30 in 7 of 8 such studies) but showed little advantage over non-intelligent tutoring software.
Strongest counter-argumentBecause it is a narrative synthesis that declines to pool effects, the headline rests on vote-counting across 28 heterogeneous quasi-experimental studies, which can mask publication bias and overweight a few high-effect comparisons.
Confidence: medium. Scored from the open-access full text including study counts, design mix, and effect-size ranges.
| ||||||
| 2026 | Watch | 70/100 | n/a |
Marcoccia, Quattrociocchi, Capraro · arXiv preprint
AI & judgment / overreliance
|
The cleanest experimental evidence yet that simply having AI advice on screen changes whether people admit they do not know something. Adults, not students, so treat it as a mechanism to design around (make "I don't know" a graded option), not a finding about your classrooms. | › |
The readFive experiments (N = 3,132; four preregistered, one direct replication) had adults answer hard questions with a real option to decline. Mere access to deliberately wrong AI advice nearly eliminated willingness to say "I don't know": participants answered more, were correct about a third as often as without AI, and their confidence nearly doubled. Accuracy incentives only partially restored suspension of judgment.
Strongest counter-argumentThe advice was engineered to be wrong on deliberately difficult trivia-style questions with incentives attached, which is not how students meet AI in schoolwork, so the near-elimination effect may overstate what happens when advice is usually right and the task is familiar.
Confidence: medium. Scored from the full abstract with design, N, preregistration status, and headline effects stated; full tables not read directly.
| ||||||
| 2026 | Watch | 67/100 | 0 |
Rismanchian, Uzun, Matayoshi, et al. · arXiv preprint (under review)
AI & learning (math)
|
A strong counterweight to AI-tutoring optimism: unsupervised generative AI access during practice may erode the learning it appears to accelerate. Useful for framing guardrails, but it is a single preprint from a vendor team pending peer review. | › |
The readA large quasi-experiment arguing that since ChatGPT, students spend 27 to 31 percent less time on AI-susceptible math problems and retain less, with the effect vanishing under proctoring, suggesting offloading rather than efficiency.
Strongest counter-argumentIdentification rests on AI susceptibility of problem type as a proxy for AI use rather than measured use, so unobserved curriculum or platform changes over 2015 to 2025 could mimic the post-ChatGPT ramp.
Confidence: medium. Scored from the full text including design, sample, effect sizes, and funding.
| ||||||
| 2026 | Watch | 66/100 | 0 |
Sungu, Lira, Duckworth (Wharton / UPenn) · SSRN preprint
AI & teacher practice
|
The first randomized field evidence on teachers using AI to prepare class materials: access alone did not lift learning, students rated the classes duller, and students of the weakest teachers lost ground. Treat "give teachers AI" as an implementation problem, not a productivity upgrade. | › |
The readA 10-week randomized field experiment (193 teachers, 2,800+ middle and high school students in a Turkish private school chain) gave treatment teachers a curriculum-customized ChatGPT assistant. Student intrinsic motivation fell 0.11 SD, classes were rated less enjoyable and less important, average achievement on externally administered exams was unchanged, and achievement and confidence declined for students of lower-performing teachers.
Strongest counter-argumentControl teachers were free to use consumer AI tools, so this estimates a customized assistant versus the status quo, not AI versus no AI; and the 'harms teaching' headline rests on a modest motivation decline plus subgroup effects, while average achievement did not change.
Confidence: medium-low. Scored from the SSRN abstract and detailed independent reporting (Hechinger Proof Points, 2026-07-13); full tables not read directly, so effect and validity dimensions were scored conservatively.
| ||||||
| 2026 | Watch | 72/100 | 6 |
Hadra, Cambridge, Mesbah · Int'l Journal for Educational Integrity
AI detection
|
Do not use AI-detector scores as standalone evidence in integrity cases. The false-classification risk is high enough that any policy should require corroborating evidence and a fair appeals process. | › |
The readTesting 192 balanced texts against Turnitin and Originality, overall accuracy was only 61 percent and 69 percent respectively, with both detectors failing badly on mixed human-AI text and degrading on longer and scientific writing.
Strongest counter-argumentThe accuracy figures rest on only 192 texts and two detectors at a single point in time, and because detector models are retrained frequently, the specific numbers may be stale soon even though the broad unreliability finding is robust.
Confidence: medium. Verified title, authors, journal, design, sample, and the no-competing-interest statement.
| ||||||
| 2025 | Watch | 70/100 | 12 |
Zhao, Yue, Sun, Jiang, Li · Journal of Intelligence
Evidence base (meta-analysis)
|
Converging evidence that generative AI can meaningfully support higher-order thinking, strongest for K-12 and 8 to 16 week interventions. High unexplained heterogeneity and no preregistration mean the pooled number is a range, not a guarantee. | › |
The readA peer-reviewed random-effects meta-analysis of 29 experiments and quasi-experiments (59 effect sizes) on generative AI and higher-order thinking. Pooled Hedges g = 0.61 (95% CI 0.49 to 0.73): moderate gains in problem-solving, critical thinking, and creativity. Heterogeneity was high.
Strongest counter-argumentAn I-squared of 77 percent means effects vary enormously, so a single pooled g masks contexts where generative AI does little or even harms; gray literature excluded and no PROSPERO registration also raise selective-reporting risk.
Confidence: high. Full article read: effect sizes, moderators, Egger's test, heterogeneity, and funding all verified.
| ||||||
| 2025 | Real | 79/100 | 2 |
Heinrich, Baily, Chen, et al. · PLOS One
AI grading / workload
|
Solid evidence that rubric-anchored AI grading of short answers performs on par with instructors and can save grading time, with the caveat that it was one university in one discipline and did not measure student learning. | › |
The readA preregistered RCT: 3,080 short-answer gradings across 271 students in four courses, with responses randomized to GPT-4-assisted versus human grading. AI grading was statistically indistinguishable from human grading; human feedback was rated only marginally more helpful (about 2.1 percentage points).
Strongest counter-argumentConcordance with human grades is not the same as accuracy; human grading is an imperfect gold standard, and the outcomes are proxies (regrade rates, perceived helpfulness) rather than learning. One institution, one short-answer format.
Confidence: high. Full peer-reviewed article read: randomization, sample, preregistration, effect sizes, and funding all verified.
| ||||||
| 2023 | Watch | 73/100 | 485 |
Weber-Wulff, Anohina-Naumeca, et al. · Int'l Journal for Educational Integrity
AI detection
|
Strong, independent, widely-cited evidence not to base disciplinary decisions on AI-detector scores. Treat the specific numbers as dated but the unreliability conclusion as robust. | › |
The readAcross roughly 14 detection tools, AI-text detectors were neither accurate nor reliable, were easily defeated by paraphrasing and translation, and skewed toward labeling text as human-written.
Strongest counter-argumentA vendor would call the findings obsolete (2023-era detectors vs GPT-3.5). The defense: the structural failure modes, paraphrasing and translation defeating detection plus a bias toward 'human,' have been independently replicated, so the qualitative conclusion is more durable than the percentages.
Confidence: medium. Scored from the abstract, metadata, and detailed secondary summaries; the full PDF was not read end to end.
| ||||||
| 2024 | Real | 80/100 | 36 |
Wang, Ribeiro, Robinson, Loeb & Demszky (Stanford) · arXiv, preregistered RCT
Human-AI tutoring at scale
|
A low-cost AI assist for existing human tutors is a credible way to lift outcomes, and it is most worth funding where your tutor bench is weakest rather than as a replacement for strong tutors. | › |
The readIn a tutor-randomized, preregistered RCT of 900 tutors and 1,800 K-12 students, giving tutors real-time AI guidance raised topic mastery by about 4 percentage points overall, with the largest gains (about 9 points) among initially lower-rated tutors.
Strongest counter-argumentThe research team is evaluating its own tool on a single tutoring platform, the work is not yet peer reviewed, and the 4 percentage point average effect is small with the real benefit concentrated almost entirely in lower-rated tutors.
Confidence: medium. Verified design, sample, effect size, and preregistration; funding scored on visible facts.
| ||||||
| 2024 | Watch | 72/100 | n/a |
NFER, for the Education Endowment Foundation · EEF/NFER evaluation report
Teacher workload
|
The cleanest causal evidence to date that AI can reduce teacher prep workload without hurting resource quality. Treat the 25-minute figure as directional given self-report, and note it says nothing about student learning. | › |
The readA school-randomized RCT of 259 science teachers across 68 English secondary schools. The ChatGPT arm cut weekly lesson-prep time from 81.5 to 56.2 minutes (about 31 percent, roughly 25 minutes) with no drop in blindly rated resource quality.
Strongest counter-argumentThe 25-minute weekly saving rests entirely on teachers' self-reported planning diaries, and teachers knew which arm they were in, so demand and recall effects could inflate the gap. It never tested whether pupils learned more.
Confidence: high. Design and results read across the EEF project page and NFER summary. Citation count: not indexed (gray-literature report).
| ||||||
| 2025 | Watch | 50/100 | n/a |
Savoldi, Attanasio, et al. · arXiv preprint
Equity & access
|
Population-representative evidence that AI access and use gaps track existing socioeconomic and gender lines, so an unmanaged AI rollout risks widening inequality. Weigh cautiously: it is a preprint, correlational, and about Italian adults generally. | › |
The readA representative quota-sampled survey (n = 1,906, stratified to the Italian population) found less-educated, older, lower-income, and female respondents adopt and use generative AI less; 40 percent cite competence barriers.
Strongest counter-argumentThe design is purely cross-sectional and correlational, so it documents divides but cannot show AI is causing or will widen inequality, and its numbers should not be transported to a classroom or another country without confirmation.
Confidence: low. Abstract and partial PDF only; some statistics not verified. Citation count: not retrievable.
| ||||||
| 2026 | Watch | 55/100 | n/a |
Chambers, Kelley · AIED 2025 (Springer LNCS 15879), via arXiv
AI detection & equity
|
First empirical test of the claim that AI detectors disproportionately flag autistic writers: flag rates were under 2 percent overall, but significantly higher for likely-autistic authors. One more independent reason not to treat a detector score as evidence against a specific student, especially one with a disability. | › |
The readA corpus study of roughly 60,000 Reddit posts split into likely-autistic and general-Reddit subcorpora, run through OpenAI's GPT-2 output detector. Under 2 percent of either subcorpus was flagged as AI-generated, but significantly more likely-autistic texts were flagged, and the textual features driving it did not map cleanly onto known features of AI text.
Strongest counter-argumentThe autism labels are proxies (subreddit membership, not diagnoses) and the detector tested is OpenAI's retired GPT-2 model, so the specific flag rates say little about the commercial detectors schools actually run today; what survives is the direction, which converges with the documented bias against non-native English writers.
Confidence: medium. Peer-reviewed venue (AIED 2025) and full abstract read; full tables not read directly.
| ||||||
| 2026 | Watch | 54/100 | n/a |
Benazet i Montobbio, Rotter & Hernández-Leo · arXiv (preprint)
AI literacy & PD design
|
Relevant to how you design AI training, not to whether AI helps students. Hands-on beat lecture immediately, then the gap closed by five weeks. If you are planning a one-off AI session for staff, this is the argument for building in a follow-up rather than expecting one afternoon to hold. | › |
The readA quasi-experiment with 126 first-year engineering undergraduates compared a two-hour hands-on session on learning with generative AI against a lecture-based version of the same content, measuring metacognitive awareness before, immediately after, and again five weeks later. The hands-on group scored higher on engagement and on knowledge of cognition immediately after. By five weeks the two groups had converged on those measures, with only the hands-on group showing a continued within-group rise in regulation of cognition. No numerical effect sizes are reported in the abstract.
Strongest counter-argumentThe between-group advantage disappears at the only timepoint a school would plan around. What is left is a within-group trend on one subscale of a self-report instrument, which is a much smaller claim than "experiential instruction works better."
Confidence: medium-low. Scored from the arXiv abstract and metadata; effect sizes are not reported there, so magnitude and statistical-validity dimensions were scored conservatively. Citation count: too new, none recorded.
| ||||||
| 2026 | Watch | 47/100 | n/a |
Nagashima, Siegrist, Scholz, et al. · Proc. ACM HCI (CSCW 2026), via arXiv
Student-AI control & trust
|
Early qualitative evidence that teachers and students want different things from classroom AI, especially on how much control students get and how much they trust it. Useful as a checklist of rollout questions to ask both groups, not as evidence for any policy. | › |
The readA storyboard speed-dating study with 16 school students and 15 teachers in Germany found systematic misalignment between teacher and student views on student-AI decision-making control, on how much each side trusts AI, and on the social and emotional side of learning with it. Accepted at CSCW 2026.
Strongest counter-argumentThirty-one participants in one country reacting to storyboards measure stated preferences, not behavior; nothing here shows what teachers or students actually do with control once they have it, and the specific misalignments may not describe any other school system.
Confidence: medium. Scored from the abstract and venue metadata (peer-reviewed CSCW acceptance confirmed); full findings not read directly. Citation count: too new, none recorded.
| ||||||
| 2024 | Watch | 63/100 | 126 |
Lee, Pope, Miles, Zarate · Computers and Education: Artificial Intelligence
Academic integrity
|
A useful counterweight to moral-panic framing, but the flat self-report likely undercounts AI use that students do not label as cheating, so do not treat it as proof that AI cheating is contained. | › |
The readSelf-reported high school cheating stayed roughly stable after ChatGPT's release, with changes varying by cheating type.
Strongest counter-argumentThe reassuring 'cheating did not rise' result may be a measurement artifact. If students do not categorize AI assistance as cheating, self-reported rates stay flat even as AI-assisted shortcutting grows.
Confidence: low. Scored from abstract and secondary summaries; full methodology and exact sample size were paywalled.
| ||||||
| 2023 | Watch | 64/100 | 60 |
Thomas, Lin, Gatz, et al. · ACM Learning Analytics & Knowledge
Human-AI tutoring
|
Encouraging directional evidence that a human-plus-AI model can lift outcomes for struggling students at a defined cost, but the quasi-experimental design means it should inform pilots, not procurement. | › |
The readAcross three urban low-income US middle schools (585 students), adding human tutors supported by AI tools to math software was associated with better proficiency and usage, with larger apparent gains for lower-achieving students, at roughly $700 per student per year.
Strongest counter-argumentWithout randomization across three differently structured deployments, the positive differences could reflect selection into the hybrid condition rather than the tutoring itself, and no quantified effect sizes are reported.
Confidence: low. Scored largely from the abstract; effect sizes could not be verified, forcing conservative scoring.
| ||||||
| 2024 | Watch | 74/100 | n/a |
Demszky, Liu, Hill, Sanghi & Chung · EdWorkingPaper 23-875
AI coaching for teachers
|
Pre-registered randomized evidence that automated feedback can shift one specific teaching practice in real classrooms. Useful for professional learning pilots, but it measures teacher talk, not student learning, so do not let a vendor cite it as an achievement result. | › |
The readA pre-registered randomized controlled trial with 224 Utah mathematics and science teachers, run in partnership with TeachFX, testing automated email feedback on "focusing questions" that press students for explanation. Teachers opened the feedback emails 53 to 65 percent of the time, and their use of focusing questions rose 20 percent (p < 0.01) against control. No other teaching practices moved. Interviews with 13 teachers surfaced skepticism about accuracy, data privacy concerns and time constraints.
Strongest counter-argumentThe outcome measured is teacher talk, a process proxy, not student learning. A 20 percent increase off a low base may be small in absolute terms, and the study cannot show that any student understood more as a result.
Confidence: medium-high. Scored from the EdWorkingPapers abstract, which states the pre-registration, sample, effect size and p-value directly. Citation count not shown because a verified count was not available at scoring time.
| ||||||
| 2026 | Watch | 72/100 | n/a |
Mata, Russell & Page · EdWorkingPaper 26-1409
Chatbots & student support
|
A rare long-horizon randomized test of a support chatbot. It moved administrative task completion and nothing else. Buy this category to reduce paperwork friction, not to move learning, and say so out loud when you buy it. | › |
The readA four-year randomized controlled trial of an AI-enabled text-messaging chatbot at a large urban public university, combined with system observation and administrator interviews. Students stayed receptive over time and impacts concentrated in improved completion of time-sensitive administrative tasks, with no detectable effects on academic performance or persistence. Centralized ownership and flexible communication emerged as conditions for sustained implementation.
Strongest counter-argumentA precise null is not proof of absence. The study may be underpowered to detect small but real effects on persistence, and a single institution with one chatbot cannot rule out that a different implementation would move learning.
Confidence: medium. Scored from the EdWorkingPapers abstract and metadata; full tables and power calculations not read directly. Citation count not shown because the paper is too recent for a stable count.
| ||||||
Every research candidate is run through the Research Integrity Gate before it earns a verdict. Nine dimensions, scored 0 to 100. Expand any one to see the exact criterion the pipeline scores against, taken verbatim from the ranking prompt.
Is this an RCT or a strong quasi-experimental design (diff-in-diff, RD, IV, matched comparison) with confounds controlled, or is it merely correlational? Reward clean causal identification; demote "X is associated with Y" framing dressed up as causal.
Would a principal, district leader, or department head actually change a decision because of this? Inert-but-interesting scores low here too.
Adequate N, representative sample, and adequately powered to detect the claimed effect. Underpowered or convenience samples score low.
Is the effect big enough to matter in a real school, not just statistically detectable? A significant-but-trivial effect scores low.
Watch for p-hacking, uncorrected multiple comparisons, and garden-of-forking-paths signals (many outcomes, flexible specifications, results that hinge on one cut of the data). Pre-specified, robust analyses score high.
Preprint scores lower than peer-reviewed; WWC-vetted scores highest. Preregistration adds points; an undisclosed analysis path loses them.
Does it replicate or converge with the existing body of evidence, or is it a lone surprising result? Reward convergence; treat outliers with caution.
Who funded and authored it? Vendor-funded or self-funded studies of the funder's own product are demoted hard. Independent funding scores high.
Does the finding transfer to a working leader's context, to real classrooms, real constraints, and real populations, or only to a lab or a single atypical site?
Then a devil's advocate tries to refute the study and names the single strongest counter-argument. Any flaw rated critical caps the score and blocks a passing verdict. A study is Real only at 75 or higher with no critical flaw, Watch when the direction is real but the evidence is not there yet, and anything weaker is held back rather than published.
How the list is ordered. The Gate score decides whether a paper appears, and it also shapes the order. Order blends three normalized signals: the Gate score (50 percent), citation velocity in citations per year (30 percent), and recency (20 percent). So a rigorous study surfaces even with fewer citations, an influential one surfaces even if older, and a strong new study is not buried under older, more-cited ones. A paper with no reliable citation count contributes zero on the citation axis and is placed by quality and recency, shown as n/a.
The Gate adapts the architecture of three open peer-review skills built for the Claude ecosystem. Full credit to their creators:
Citation counts via Semantic Scholar / OpenAlex, retrieved 2026-06-29.