Research

Every study here is run through one rigorous bar before it earns a verdict. Ordered by a blend of how well it holds up, how influential it is, and how recent. Click any row for the full scoring.

The bar every study clears

Each paper is scored 0 to 100 across nine dimensions of research quality (causal design and decision-relevance carry the most weight), then a devil's advocate tries to refute it. Any flaw rated critical caps the score and blocks a passing verdict outright. A study is REAL only at 75 or higher with no critical flaw, WATCH if the direction is real but the evidence is not there yet, and anything weaker is held back. The order blends three normalized signals: the Gate score, citation velocity (citations per year), and recency. A rigorous study surfaces even with fewer citations, an influential one even if older, and a strong new study is not buried.

Real · 75+, no critical flaw Watch · promising, evidence not there yet
Scored studies · ordered by influence + recency
YearVerdictScoreCitedStudyWhat it means for leaders
2026 Real 85/100 n/a
Allcott, Baron, Dee, Duckworth, Gentzkow & Jacob · NBER Working Paper 35132
Phones, attention & achievement
The strongest causal evidence yet that phone restriction changes behaviour and climate but not test scores. Budget and justify a ban on attention and climate over a multi-year horizon, and plan for a worse first year, not on promised achievement gains.
Causal design13/15
Decision-relevance15/15
Sample & power10/10
Effect size8/10
Stat validity8/10
Peer-review5/10
Replication9/10
Conflicts / funding9/10
Generalizability9/10
The readA staggered difference-in-differences study of lockable phone pouches nationwide, combining large-scale surveys, GPS pings, standardized test scores, school administrative records and the largest pouch provider's sales records. Pouch adoption substantially reduces phone use. In year one disciplinary incidents rise and student well-being falls; both reverse in later years. Average effects on test scores are consistently close to zero, with modest positive high-school effects (particularly math) and small negative middle-school effects.
Strongest counter-argumentPouch adoption is not randomly assigned, and staggered difference-in-differences can be biased when treatment effects vary across adoption cohorts. The GPS compliance evidence, however, rules out the strongest rival explanation, that the bans were simply never enforced.
  • MinorNBER working paper; not yet peer reviewed.
  • MinorMeasures one mechanism, lockable pouches, so it does not settle what a bell-to-bell ban enforced another way would produce.
  • MinorMany outcomes examined, creating some multiple-comparison exposure, though nulls are reported evenhandedly.
Confidence: high. Scored from the NBER abstract and paper listing; design, data sources and directional findings are stated explicitly by the authors. Citation count not shown because the paper is too recent for a stable count.
2024 Real 76/100 513
Fan, Tang, Le, et al. · British Journal of Educational Technology
AI, writing & critical thinking
Credible randomized evidence that AI help improves the artifact more than the learner. Design assessments and tooling so AI assists the process rather than substituting for it.
Causal design13/15
Decision-relevance13/15
Sample & power7/10
Effect size7/10
Stat validity8/10
Peer-review10/10
Replication5/10
Conflicts / funding8/10
Generalizability5/10
The readIn a randomized four-condition experiment, ChatGPT users improved their essay scores the most but gained no more knowledge transfer than other groups, suggesting AI support can boost the output while offloading the thinking.
Strongest counter-argumentThe 'metacognitive laziness' label over-reads a null. ChatGPT users produced better essays and merely showed no advantage on knowledge transfer in one short lab task, and a missing transfer benefit in an underpowered single session is not evidence of cognitive harm.
  • MajorSmall cells: 117 students across four arms (~29 per group), so the null knowledge-transfer finding is likely underpowered.
  • MinorSingle lab essay task, one session, 2023-era ChatGPT; limited ecological validity.
  • MinorNo preregistration or direct replication.
Confidence: medium. Scored from the abstract and detailed secondary summaries confirming the four-arm design and N=117; full tables not read directly.
2025 Watch 51/100 408
Kosmyna, Hauptmann, et al. (MIT Media Lab) · arXiv preprint
AI, writing & critical thinking
Provocative and directionally consistent with offloading concerns, but too small and unreviewed to act on. Treat it as a hypothesis to watch, not proof that AI assistance damages thinking.
Causal design10/15
Decision-relevance11/15
Sample & power4/10
Effect size5/10
Stat validity4/10
Peer-review2/10
Replication3/10
Conflicts / funding8/10
Generalizability4/10
The readAcross 54 participants writing essays under EEG, LLM users showed the weakest brain connectivity and the poorest recall of their own essays, which the authors frame as accumulating 'cognitive debt.'
Strongest counter-argumentThe viral 'your brain shuts off with ChatGPT' conclusion is statistically fragile: the headline crossover result rests on 18 participants in a non-randomized fourth session, with dozens of EEG comparisons, no peer review, and a single artificial essay task.
  • MajorNot peer reviewed at release (arXiv preprint), a status the authors themselves emphasized.
  • MajorHeadline 'cognitive debt' claim relies on 18 participants in session 4, non-randomized, heavy multiple-comparison EEG analysis.
  • MinorSingle SAT-style essay task and a Boston-area convenience sample; widely over-interpreted in media.
Confidence: low. Scored from the abstract and detailed secondary summaries; full statistical detail not read directly.
2025 Watch 76/100 98
Kestin, Miller, Klales, et al. · Scientific Reports
AI tutoring vs active learning
A well-engineered AI tutor can match or beat good instruction in a controlled lesson, but the evidence is too narrow to justify replacing classroom teaching. Treat it as promising for piloting supplemental practice, not a deployment mandate.
Causal design13/15
Decision-relevance13/15
Sample & power6/10
Effect size9/10
Stat validity8/10
Peer-review7/10
Replication6/10
Conflicts / funding9/10
Generalizability5/10
The readA crossover RCT of 194 Harvard physics students found a custom GPT-4 tutor with engineered pedagogical prompts produced post-test gains of about 0.63 SD, up to 0.73 to 1.3 SD by quantile, highly significant, in less instructional time than in-class active learning.
Strongest counter-argumentThe result comes from 194 elite Harvard physics students on just two topics over roughly 50 minutes, so novelty effects, ceiling effects, and the narrow setting make it unsafe to assume the same gains in a typical K-12 or community-college classroom.
  • MajorSingle elite institution, single subject, two short topics; weak external validity.
  • MinorNot preregistered.
  • MinorCeiling effects addressed via quantile regression but still present.
Confidence: high. Read the full published paper including methods, effect sizes, and the no-competing-interests statement.
2025 Watch 72/100 128
Letourneau, Deslandes Martineau, et al. · npj Science of Learning
Evidence base (review)
The best single map of where ITS evidence is strong versus thin in K-12. The strategic nuance for a buyer: the gain comes from structured tutoring, not necessarily the AI layer, since ITS rarely beat simpler tutoring software.
Causal design6/15
Decision-relevance13/15
Sample & power8/10
Effect size6/10
Stat validity6/10
Peer-review8/10
Replication8/10
Conflicts / funding9/10
Generalizability8/10
The readA peer-reviewed synthesis of 28 studies (4,597 K-12 students) finding intelligent tutoring systems generally beat human-teacher-only conditions (Hedges g roughly 0.68 to 1.30 in 7 of 8 such studies) but showed little advantage over non-intelligent tutoring software.
Strongest counter-argumentBecause it is a narrative synthesis that declines to pool effects, the headline rests on vote-counting across 28 heterogeneous quasi-experimental studies, which can mask publication bias and overweight a few high-effect comparisons.
  • MajorNarrative review, not a meta-analysis; no pooled effect estimate, so synthesis relies on vote-counting.
  • MinorUnderlying evidence base is overwhelmingly quasi-experimental.
Confidence: medium. Scored from the open-access full text including study counts, design mix, and effect-size ranges.
2026 Watch 70/100 n/a
Marcoccia, Quattrociocchi, Capraro · arXiv preprint
AI & judgment / overreliance
The cleanest experimental evidence yet that simply having AI advice on screen changes whether people admit they do not know something. Adults, not students, so treat it as a mechanism to design around (make "I don't know" a graded option), not a finding about your classrooms.
Causal design13/15
Decision-relevance9/15
Sample & power8/10
Effect size8/10
Stat validity9/10
Peer-review3/10
Replication7/10
Conflicts / funding9/10
Generalizability4/10
The readFive experiments (N = 3,132; four preregistered, one direct replication) had adults answer hard questions with a real option to decline. Mere access to deliberately wrong AI advice nearly eliminated willingness to say "I don't know": participants answered more, were correct about a third as often as without AI, and their confidence nearly doubled. Accuracy incentives only partially restored suspension of judgment.
Strongest counter-argumentThe advice was engineered to be wrong on deliberately difficult trivia-style questions with incentives attached, which is not how students meet AI in schoolwork, so the near-elimination effect may overstate what happens when advice is usually right and the task is familiar.
  • MajorAdult online participants and artificial question sets; no student sample, so school generalization is untested.
  • MajorNot peer reviewed (arXiv preprint, July 2026).
  • MinorDeliberately wrong advice isolates the mechanism cleanly but caps ecological validity.
Confidence: medium. Scored from the full abstract with design, N, preregistration status, and headline effects stated; full tables not read directly.
2026 Watch 67/100 0
Rismanchian, Uzun, Matayoshi, et al. · arXiv preprint (under review)
AI & learning (math)
A strong counterweight to AI-tutoring optimism: unsupervised generative AI access during practice may erode the learning it appears to accelerate. Useful for framing guardrails, but it is a single preprint from a vendor team pending peer review.
Causal design10/15
Decision-relevance14/15
Sample & power10/10
Effect size9/10
Stat validity8/10
Peer-review2/10
Replication4/10
Conflicts / funding4/10
Generalizability6/10
The readA large quasi-experiment arguing that since ChatGPT, students spend 27 to 31 percent less time on AI-susceptible math problems and retain less, with the effect vanishing under proctoring, suggesting offloading rather than efficiency.
Strongest counter-argumentIdentification rests on AI susceptibility of problem type as a proxy for AI use rather than measured use, so unobserved curriculum or platform changes over 2015 to 2025 could mimic the post-ChatGPT ramp.
  • MajorFour of five authors are employed by McGraw Hill, owner of the ALEKS platform studied.
  • MajorarXiv preprint, not yet peer reviewed.
  • MinorNot preregistered; AI use inferred from a proxy, never directly observed.
Confidence: medium. Scored from the full text including design, sample, effect sizes, and funding.
2026 Watch 66/100 0
Sungu, Lira, Duckworth (Wharton / UPenn) · SSRN preprint
AI & teacher practice
The first randomized field evidence on teachers using AI to prepare class materials: access alone did not lift learning, students rated the classes duller, and students of the weakest teachers lost ground. Treat "give teachers AI" as an implementation problem, not a productivity upgrade.
Causal design12/15
Decision-relevance13/15
Sample & power7/10
Effect size5/10
Stat validity7/10
Peer-review3/10
Replication6/10
Conflicts / funding8/10
Generalizability5/10
The readA 10-week randomized field experiment (193 teachers, 2,800+ middle and high school students in a Turkish private school chain) gave treatment teachers a curriculum-customized ChatGPT assistant. Student intrinsic motivation fell 0.11 SD, classes were rated less enjoyable and less important, average achievement on externally administered exams was unchanged, and achievement and confidence declined for students of lower-performing teachers.
Strongest counter-argumentControl teachers were free to use consumer AI tools, so this estimates a customized assistant versus the status quo, not AI versus no AI; and the 'harms teaching' headline rests on a modest motivation decline plus subgroup effects, while average achievement did not change.
  • MajorSSRN preprint (June 2026), not yet peer reviewed.
  • MajorHeadline harm is concentrated in subgroup analyses (lower-performing teachers, prior heavy AI users); multiple outcomes raise forking-paths risk.
  • MinorSingle private school network in one country; a customized tool, so transfer to US districts is untested.
  • MinorMechanism unobserved: no classroom observation or analysis of the AI-generated materials.
Confidence: medium-low. Scored from the SSRN abstract and detailed independent reporting (Hechinger Proof Points, 2026-07-13); full tables not read directly, so effect and validity dimensions were scored conservatively.
2026 Watch 72/100 6
Hadra, Cambridge, Mesbah · Int'l Journal for Educational Integrity
AI detection
Do not use AI-detector scores as standalone evidence in integrity cases. The false-classification risk is high enough that any policy should require corroborating evidence and a fair appeals process.
Causal design9/15
Decision-relevance13/15
Sample & power5/10
Effect size7/10
Stat validity7/10
Peer-review8/10
Replication8/10
Conflicts / funding9/10
Generalizability6/10
The readTesting 192 balanced texts against Turnitin and Originality, overall accuracy was only 61 percent and 69 percent respectively, with both detectors failing badly on mixed human-AI text and degrading on longer and scientific writing.
Strongest counter-argumentThe accuracy figures rest on only 192 texts and two detectors at a single point in time, and because detector models are retrained frequently, the specific numbers may be stale soon even though the broad unreliability finding is robust.
  • MajorSmall corpus (192 texts) and only two commercial detectors tested.
  • MinorDetector versions evolve quickly, limiting the shelf life of exact accuracy numbers.
Confidence: medium. Verified title, authors, journal, design, sample, and the no-competing-interest statement.
2025 Watch 70/100 12
Zhao, Yue, Sun, Jiang, Li · Journal of Intelligence
Evidence base (meta-analysis)
Converging evidence that generative AI can meaningfully support higher-order thinking, strongest for K-12 and 8 to 16 week interventions. High unexplained heterogeneity and no preregistration mean the pooled number is a range, not a guarantee.
Causal design10/15
Decision-relevance12/15
Sample & power7/10
Effect size8/10
Stat validity8/10
Peer-review6/10
Replication5/10
Conflicts / funding8/10
Generalizability6/10
The readA peer-reviewed random-effects meta-analysis of 29 experiments and quasi-experiments (59 effect sizes) on generative AI and higher-order thinking. Pooled Hedges g = 0.61 (95% CI 0.49 to 0.73): moderate gains in problem-solving, critical thinking, and creativity. Heterogeneity was high.
Strongest counter-argumentAn I-squared of 77 percent means effects vary enormously, so a single pooled g masks contexts where generative AI does little or even harms; gray literature excluded and no PROSPERO registration also raise selective-reporting risk.
  • MajorHigh unexplained heterogeneity (I-squared = 77 percent) limits interpretability of the pooled effect.
  • MajorNo preregistration and gray literature excluded, raising publication and selection bias risk.
  • MinorIncludes quasi-experiments; weighted toward short Chinese and English-language studies.
Confidence: high. Full article read: effect sizes, moderators, Egger's test, heterogeneity, and funding all verified.
2025 Real 79/100 2
Heinrich, Baily, Chen, et al. · PLOS One
AI grading / workload
Solid evidence that rubric-anchored AI grading of short answers performs on par with instructors and can save grading time, with the caveat that it was one university in one discipline and did not measure student learning.
Causal design14/15
Decision-relevance12/15
Sample & power8/10
Effect size8/10
Stat validity9/10
Peer-review9/10
Replication5/10
Conflicts / funding9/10
Generalizability5/10
The readA preregistered RCT: 3,080 short-answer gradings across 271 students in four courses, with responses randomized to GPT-4-assisted versus human grading. AI grading was statistically indistinguishable from human grading; human feedback was rated only marginally more helpful (about 2.1 percentage points).
Strongest counter-argumentConcordance with human grades is not the same as accuracy; human grading is an imperfect gold standard, and the outcomes are proxies (regrade rates, perceived helpfulness) rather than learning. One institution, one short-answer format.
  • MinorSingle institution and one discipline limit generalizability.
  • MinorOutcomes are grading-concordance and perception proxies, not student learning.
Confidence: high. Full peer-reviewed article read: randomization, sample, preregistration, effect sizes, and funding all verified.
2023 Watch 73/100 485
Weber-Wulff, Anohina-Naumeca, et al. · Int'l Journal for Educational Integrity
AI detection
Strong, independent, widely-cited evidence not to base disciplinary decisions on AI-detector scores. Treat the specific numbers as dated but the unreliability conclusion as robust.
Causal design4/15
Decision-relevance14/15
Sample & power7/10
Effect size8/10
Stat validity7/10
Peer-review9/10
Replication9/10
Conflicts / funding9/10
Generalizability6/10
The readAcross roughly 14 detection tools, AI-text detectors were neither accurate nor reliable, were easily defeated by paraphrasing and translation, and skewed toward labeling text as human-written.
Strongest counter-argumentA vendor would call the findings obsolete (2023-era detectors vs GPT-3.5). The defense: the structural failure modes, paraphrasing and translation defeating detection plus a bias toward 'human,' have been independently replicated, so the qualitative conclusion is more durable than the percentages.
  • MajorRapid tool obsolescence: the detectors and models tested are now superseded.
  • MinorLargely descriptive classification scoring rather than formal diagnostic-accuracy statistics.
  • MinorModest corpus mixing human, AI, machine-translated, and obfuscated text.
Confidence: medium. Scored from the abstract, metadata, and detailed secondary summaries; the full PDF was not read end to end.
2024 Real 80/100 36
Wang, Ribeiro, Robinson, Loeb & Demszky (Stanford) · arXiv, preregistered RCT
Human-AI tutoring at scale
A low-cost AI assist for existing human tutors is a credible way to lift outcomes, and it is most worth funding where your tutor bench is weakest rather than as a replacement for strong tutors.
Causal design13/15
Decision-relevance14/15
Sample & power9/10
Effect size7/10
Stat validity8/10
Peer-review6/10
Replication6/10
Conflicts / funding8/10
Generalizability9/10
The readIn a tutor-randomized, preregistered RCT of 900 tutors and 1,800 K-12 students, giving tutors real-time AI guidance raised topic mastery by about 4 percentage points overall, with the largest gains (about 9 points) among initially lower-rated tutors.
Strongest counter-argumentThe research team is evaluating its own tool on a single tutoring platform, the work is not yet peer reviewed, and the 4 percentage point average effect is small with the real benefit concentrated almost entirely in lower-rated tutors.
  • MinorarXiv preprint, not yet peer reviewed (mitigated by OSF preregistration).
  • MinorCreators evaluating their own system; single tutoring platform context.
Confidence: medium. Verified design, sample, effect size, and preregistration; funding scored on visible facts.
2024 Watch 72/100 n/a
NFER, for the Education Endowment Foundation · EEF/NFER evaluation report
Teacher workload
The cleanest causal evidence to date that AI can reduce teacher prep workload without hurting resource quality. Treat the 25-minute figure as directional given self-report, and note it says nothing about student learning.
Causal design13/15
Decision-relevance14/15
Sample & power7/10
Effect size7/10
Stat validity7/10
Peer-review6/10
Replication4/10
Conflicts / funding8/10
Generalizability6/10
The readA school-randomized RCT of 259 science teachers across 68 English secondary schools. The ChatGPT arm cut weekly lesson-prep time from 81.5 to 56.2 minutes (about 31 percent, roughly 25 minutes) with no drop in blindly rated resource quality.
Strongest counter-argumentThe 25-minute weekly saving rests entirely on teachers' self-reported planning diaries, and teachers knew which arm they were in, so demand and recall effects could inflate the gap. It never tested whether pupils learned more.
  • MajorPrimary outcome (planning time) is self-reported by unblinded teachers, vulnerable to demand and recall bias.
  • MinorPublished as an EEF/NFER evaluation report, not a peer-reviewed journal article.
  • MinorSingle context (England KS3 science, one term); no pupil learning outcome measured.
Confidence: high. Design and results read across the EEF project page and NFER summary. Citation count: not indexed (gray-literature report).
2025 Watch 50/100 n/a
Savoldi, Attanasio, et al. · arXiv preprint
Equity & access
Population-representative evidence that AI access and use gaps track existing socioeconomic and gender lines, so an unmanaged AI rollout risks widening inequality. Weigh cautiously: it is a preprint, correlational, and about Italian adults generally.
Causal design4/15
Decision-relevance10/15
Sample & power8/10
Effect size5/10
Stat validity6/10
Peer-review2/10
Replication4/10
Conflicts / funding6/10
Generalizability5/10
The readA representative quota-sampled survey (n = 1,906, stratified to the Italian population) found less-educated, older, lower-income, and female respondents adopt and use generative AI less; 40 percent cite competence barriers.
Strongest counter-argumentThe design is purely cross-sectional and correlational, so it documents divides but cannot show AI is causing or will widen inequality, and its numbers should not be transported to a classroom or another country without confirmation.
  • MajorNot peer reviewed (arXiv preprint).
  • MajorPurely correlational and cross-sectional; no causal identification.
  • MinorGeneral Italian adult population rather than students; transferability limited.
Confidence: low. Abstract and partial PDF only; some statistics not verified. Citation count: not retrievable.
2026 Watch 55/100 n/a
Chambers, Kelley · AIED 2025 (Springer LNCS 15879), via arXiv
AI detection & equity
First empirical test of the claim that AI detectors disproportionately flag autistic writers: flag rates were under 2 percent overall, but significantly higher for likely-autistic authors. One more independent reason not to treat a detector score as evidence against a specific student, especially one with a disability.
Causal design5/15
Decision-relevance9/15
Sample & power5/10
Effect size4/10
Stat validity6/10
Peer-review7/10
Replication7/10
Conflicts / funding9/10
Generalizability3/10
The readA corpus study of roughly 60,000 Reddit posts split into likely-autistic and general-Reddit subcorpora, run through OpenAI's GPT-2 output detector. Under 2 percent of either subcorpus was flagged as AI-generated, but significantly more likely-autistic texts were flagged, and the textual features driving it did not map cleanly onto known features of AI text.
Strongest counter-argumentThe autism labels are proxies (subreddit membership, not diagnoses) and the detector tested is OpenAI's retired GPT-2 model, so the specific flag rates say little about the commercial detectors schools actually run today; what survives is the direction, which converges with the documented bias against non-native English writers.
  • MajorObsolete detector: OpenAI's GPT-2 output detector, not the commercial tools in current school use.
  • MajorProxy labeling: "likely-autistic" inferred from subreddit membership, not verified authorship.
  • MinorReddit posts, not student schoolwork; absolute flag rates were low in both groups.
Confidence: medium. Peer-reviewed venue (AIED 2025) and full abstract read; full tables not read directly.
2026 Watch 54/100 n/a
Benazet i Montobbio, Rotter & Hernández-Leo · arXiv (preprint)
AI literacy & PD design
Relevant to how you design AI training, not to whether AI helps students. Hands-on beat lecture immediately, then the gap closed by five weeks. If you are planning a one-off AI session for staff, this is the argument for building in a follow-up rather than expecting one afternoon to hold.
Causal design9/15
Decision-relevance9/15
Sample & power5/10
Effect size3/10
Stat validity6/10
Peer-review3/10
Replication6/10
Conflicts / funding9/10
Generalizability4/10
The readA quasi-experiment with 126 first-year engineering undergraduates compared a two-hour hands-on session on learning with generative AI against a lecture-based version of the same content, measuring metacognitive awareness before, immediately after, and again five weeks later. The hands-on group scored higher on engagement and on knowledge of cognition immediately after. By five weeks the two groups had converged on those measures, with only the hands-on group showing a continued within-group rise in regulation of cognition. No numerical effect sizes are reported in the abstract.
Strongest counter-argumentThe between-group advantage disappears at the only timepoint a school would plan around. What is left is a within-group trend on one subscale of a self-report instrument, which is a much smaller claim than "experiential instruction works better."
  • MajorBetween-group difference did not survive to the five-week follow-up; the surviving result is a within-group trend on a single subscale.
  • MinorSelf-reported metacognitive awareness, not observed behavior or learning outcomes.
  • MinorPreprint, no preregistration stated; 126 students in one first-year engineering course in one institution.
  • MinorHigher education, not K-12; transfer to school staff PD is plausible but untested.
Confidence: medium-low. Scored from the arXiv abstract and metadata; effect sizes are not reported there, so magnitude and statistical-validity dimensions were scored conservatively. Citation count: too new, none recorded.
2026 Watch 47/100 n/a
Nagashima, Siegrist, Scholz, et al. · Proc. ACM HCI (CSCW 2026), via arXiv
Student-AI control & trust
Early qualitative evidence that teachers and students want different things from classroom AI, especially on how much control students get and how much they trust it. Useful as a checklist of rollout questions to ask both groups, not as evidence for any policy.
Causal design6/15
Decision-relevance8/15
Sample & power2/10
Effect size3/10
Stat validity5/10
Peer-review8/10
Replication3/10
Conflicts / funding9/10
Generalizability3/10
The readA storyboard speed-dating study with 16 school students and 15 teachers in Germany found systematic misalignment between teacher and student views on student-AI decision-making control, on how much each side trusts AI, and on the social and emotional side of learning with it. Accepted at CSCW 2026.
Strongest counter-argumentThirty-one participants in one country reacting to storyboards measure stated preferences, not behavior; nothing here shows what teachers or students actually do with control once they have it, and the specific misalignments may not describe any other school system.
  • MajorVery small sample (16 students, 15 teachers) in a single national context.
  • MinorStoryboard speed-dating elicits stated preferences rather than observed classroom behavior.
  • MinorNo replication; first study of its pairing design in K-12 AI.
Confidence: medium. Scored from the abstract and venue metadata (peer-reviewed CSCW acceptance confirmed); full findings not read directly. Citation count: too new, none recorded.
2024 Watch 63/100 126
Lee, Pope, Miles, Zarate · Computers and Education: Artificial Intelligence
Academic integrity
A useful counterweight to moral-panic framing, but the flat self-report likely undercounts AI use that students do not label as cheating, so do not treat it as proof that AI cheating is contained.
Causal design5/15
Decision-relevance13/15
Sample & power6/10
Effect size6/10
Stat validity6/10
Peer-review9/10
Replication6/10
Conflicts / funding7/10
Generalizability5/10
The readSelf-reported high school cheating stayed roughly stable after ChatGPT's release, with changes varying by cheating type.
Strongest counter-argumentThe reassuring 'cheating did not rise' result may be a measurement artifact. If students do not categorize AI assistance as cheating, self-reported rates stay flat even as AI-assisted shortcutting grows.
  • MajorSelf-reported cheating via anonymous repeated cross-sectional surveys, vulnerable to social-desirability and shifting norms.
  • MajorOnly three Challenge Success partner schools that skew affluent and high-achieving.
  • MinorPre/post window straddles COVID-era disruption, a confound.
  • MinorStability could reflect definitional drift rather than unchanged behavior.
Confidence: low. Scored from abstract and secondary summaries; full methodology and exact sample size were paywalled.
2023 Watch 64/100 60
Thomas, Lin, Gatz, et al. · ACM Learning Analytics & Knowledge
Human-AI tutoring
Encouraging directional evidence that a human-plus-AI model can lift outcomes for struggling students at a defined cost, but the quasi-experimental design means it should inform pilots, not procurement.
Causal design8/15
Decision-relevance12/15
Sample & power7/10
Effect size5/10
Stat validity6/10
Peer-review7/10
Replication7/10
Conflicts / funding6/10
Generalizability6/10
The readAcross three urban low-income US middle schools (585 students), adding human tutors supported by AI tools to math software was associated with better proficiency and usage, with larger apparent gains for lower-achieving students, at roughly $700 per student per year.
Strongest counter-argumentWithout randomization across three differently structured deployments, the positive differences could reflect selection into the hybrid condition rather than the tutoring itself, and no quantified effect sizes are reported.
  • MajorQuasi-experimental with no randomization; conditions differ across schools.
  • MinorNo numerical effect sizes available in the abstract.
  • MinorFunding and conflict-of-interest disclosures not visible.
Confidence: low. Scored largely from the abstract; effect sizes could not be verified, forcing conservative scoring.
2024 Watch 74/100 n/a
Demszky, Liu, Hill, Sanghi & Chung · EdWorkingPaper 23-875
AI coaching for teachers
Pre-registered randomized evidence that automated feedback can shift one specific teaching practice in real classrooms. Useful for professional learning pilots, but it measures teacher talk, not student learning, so do not let a vendor cite it as an achievement result.
Causal design14/15
Decision-relevance12/15
Sample & power7/10
Effect size6/10
Stat validity9/10
Peer-review7/10
Replication8/10
Conflicts / funding5/10
Generalizability6/10
The readA pre-registered randomized controlled trial with 224 Utah mathematics and science teachers, run in partnership with TeachFX, testing automated email feedback on "focusing questions" that press students for explanation. Teachers opened the feedback emails 53 to 65 percent of the time, and their use of focusing questions rose 20 percent (p < 0.01) against control. No other teaching practices moved. Interviews with 13 teachers surfaced skepticism about accuracy, data privacy concerns and time constraints.
Strongest counter-argumentThe outcome measured is teacher talk, a process proxy, not student learning. A 20 percent increase off a low base may be small in absolute terms, and the study cannot show that any student understood more as a result.
  • MajorProcess outcome only; no student learning measured.
  • MajorConducted in partnership with TeachFX, the vendor whose product is evaluated. Disclosed, but a live conflict.
  • MinorSingle state, mathematics and science only; working paper, not peer reviewed.
Confidence: medium-high. Scored from the EdWorkingPapers abstract, which states the pre-registration, sample, effect size and p-value directly. Citation count not shown because a verified count was not available at scoring time.
2026 Watch 72/100 n/a
Mata, Russell & Page · EdWorkingPaper 26-1409
Chatbots & student support
A rare long-horizon randomized test of a support chatbot. It moved administrative task completion and nothing else. Buy this category to reduce paperwork friction, not to move learning, and say so out loud when you buy it.
Causal design13/15
Decision-relevance11/15
Sample & power8/10
Effect size6/10
Stat validity8/10
Peer-review5/10
Replication9/10
Conflicts / funding8/10
Generalizability4/10
The readA four-year randomized controlled trial of an AI-enabled text-messaging chatbot at a large urban public university, combined with system observation and administrator interviews. Students stayed receptive over time and impacts concentrated in improved completion of time-sensitive administrative tasks, with no detectable effects on academic performance or persistence. Centralized ownership and flexible communication emerged as conditions for sustained implementation.
Strongest counter-argumentA precise null is not proof of absence. The study may be underpowered to detect small but real effects on persistence, and a single institution with one chatbot cannot rule out that a different implementation would move learning.
  • MajorSingle institution and post-secondary only; transfer to K-12 is an inference, not a finding.
  • MinorWorking paper; not yet peer reviewed.
  • MinorNo preregistration stated in the abstract.
Confidence: medium. Scored from the EdWorkingPapers abstract and metadata; full tables and power calculations not read directly. Citation count not shown because the paper is too recent for a stable count.
How a study is scored

Every research candidate is run through the Research Integrity Gate before it earns a verdict. Nine dimensions, scored 0 to 100. Expand any one to see the exact criterion the pipeline scores against, taken verbatim from the ranking prompt.

Causal identification / design15

Is this an RCT or a strong quasi-experimental design (diff-in-diff, RD, IV, matched comparison) with confounds controlled, or is it merely correlational? Reward clean causal identification; demote "X is associated with Y" framing dressed up as causal.

Decision-relevance to a school / education leader15

Would a principal, district leader, or department head actually change a decision because of this? Inert-but-interesting scores low here too.

Sample & power10

Adequate N, representative sample, and adequately powered to detect the claimed effect. Underpowered or convenience samples score low.

Effect size & practical magnitude10

Is the effect big enough to matter in a real school, not just statistically detectable? A significant-but-trivial effect scores low.

Statistical validity10

Watch for p-hacking, uncorrected multiple comparisons, and garden-of-forking-paths signals (many outcomes, flexible specifications, results that hinge on one cut of the data). Pre-specified, robust analyses score high.

Peer-review / preregistration status10

Preprint scores lower than peer-reviewed; WWC-vetted scores highest. Preregistration adds points; an undisclosed analysis path loses them.

Replication & convergence with prior evidence10

Does it replicate or converge with the existing body of evidence, or is it a lone surprising result? Reward convergence; treat outliers with caution.

Conflicts of interest / funding10

Who funded and authored it? Vendor-funded or self-funded studies of the funder's own product are demoted hard. Independent funding scores high.

Generalizability to real classrooms / districts10

Does the finding transfer to a working leader's context, to real classrooms, real constraints, and real populations, or only to a lab or a single atypical site?

Then a devil's advocate tries to refute the study and names the single strongest counter-argument. Any flaw rated critical caps the score and blocks a passing verdict. A study is Real only at 75 or higher with no critical flaw, Watch when the direction is real but the evidence is not there yet, and anything weaker is held back rather than published.

How the list is ordered. The Gate score decides whether a paper appears, and it also shapes the order. Order blends three normalized signals: the Gate score (50 percent), citation velocity in citations per year (30 percent), and recency (20 percent). So a rigorous study surfaces even with fewer citations, an influential one surfaces even if older, and a strong new study is not buried under older, more-cited ones. A paper with no reliable citation count contributes zero on the citation axis and is placed by quality and recency, shown as n/a.

Attribution

The Gate adapts the architecture of three open peer-review skills built for the Claude ecosystem. Full credit to their creators:

Citation counts via Semantic Scholar / OpenAlex, retrieved 2026-06-29.

The Margin · A weekly AI-in-education digest for education leaders
Home · Methodology · My Planning Partner · @myplanningpartner