What's real in AI for your classrooms this week, and what's just loud. A weekly digest of the most practical findings and headlines for the K-12 education leader.
A school-randomised controlled trial of 259 Year 7 and Year 8 science teachers across 68 English secondary schools, run over ten weeks in the summer term of 2024, funded by the Education Endowment Foundation and the Hg Foundation and independently evaluated by the National Foundation for Educational Research. Teachers in the ChatGPT group were asked to use it to prepare lessons and resources and were given an online guide; the comparison group was asked not to use generative AI for those lessons. Lesson and resource planning time was 56.2 minutes per week in the ChatGPT group against 81.5 in the comparison group, a saving of 25.3 minutes, about a 31 percent reduction. On quality the evaluators report: "We found no evidence to suggest that the quality of the lesson resources used by the two groups differed, from an expert panel review (who did not know how the resources had been created)." The limitation that matters most is what was never on the table. The trial did not measure student learning at all. It measured preparation time and resource quality, not achievement, engagement or outcomes. And it is one subject, two year groups and one ten-week window. Twenty-five minutes a week is a genuine gift to a science teacher and it is not a transformation. When a vendor quotes you hours saved, the follow-up is not whether to believe it. It is what the time was saved from, and what happened to the students. NFER → Full report →
Disclosure, and it is the one I always make: AI for lesson preparation is exactly my category. I am building My Planning Partner in that lane, so read my take knowing I have a horse in this race.
The Harvard Crimson surveyed Harvard Faculty of Arts and Sciences professors, and Inside Higher Ed reported the results on September 17, 2026. Nearly two thirds of Harvard faculty said artificial intelligence has had a "somewhat negative" or "very negative" effect on their courses. Last year 42 percent said the same. Policy has hardened alongside the mood: only 3.5 percent said they "didn't have any such policy," down from 10 percent the year prior, only 4 percent of professors "entirely permit" AI use in class, down from 8 percent last year, and about a quarter prohibit all AI use in class, up from 20 percent. The number worth holding is the gap between suspicion and certainty. Nearly nine in 10 faculty members reported receiving student work that they "knew or believed was produced using AI," while only 64 percent reported feeling "somewhat or very confident" in their ability to distinguish between AI-generated and student-created work. Only 12 percent referred a case to the honor council. Where the claim breaks: this is a student-newspaper survey of one elite institution's faculty, the response rate, sampling method and fielding dates were not disclosed in the report, and every figure is self-report about perception rather than a measurement of how much AI writing was actually submitted. Strip out the Harvard part and this is a picture of what happens when belief outruns proof. Almost everyone thinks they have seen it. A third will tell you they cannot reliably identify it. And almost nobody files a case, which is the honest tell. The useful question at your next department meeting is not how to catch it. It is what you are willing to act on, and what happens to a student when you are wrong. Inside Higher Ed →
Calvin Isley, Johann D. Gaebler and Sharad Goel, an arXiv preprint posted September 18, 2026, analyzing roughly 7,500 applications submitted between 2020 and 2025 to a large United States public policy master's program that prohibited AI use. The authors report that most applicants in the 2025 cycle submitted an essay that was primarily AI-generated despite the prohibition. Using the November 2022 launch of ChatGPT as a natural break, they find AI availability improved essay writing quality, and yet applicants whose essays were AI-written were admitted less often than otherwise comparable applicants who did not use it. A separate follow-up experiment found admissions officers could often identify AI-written essays and rated them lower. Research Integrity Gate 59 out of 100, Watch. The major flags, stated here rather than buried: the design is a before-and-after comparison around a single date, so anything else that changed in this program's applicant pool between 2020 and 2025 is confounded with AI availability, the detection of AI authorship is itself an estimate rather than ground truth, it is one graduate program in one field, and it is a preprint that is not peer reviewed and not preregistered. This is the clearest version yet of a thing worth telling a junior directly. The essay got better and the applicant did worse. It is one program, so do not turn it into a rule. But when a student asks whether to run their personal statement through a chatbot, you now have something better to say than that it is against the rules. arXiv preprint →
On September 16, 2026, a week after the signed AFT and UFT agreement, Justin Spelhaug, President of Microsoft Elevate, published "Microsoft's commitment for AI in education." It restates the privacy standard and gives no commitment dates. Around that sit five numbers, and not one of them carries a study, a sample size or a method. Brisbane Catholic Education is credited with a "275% increase in learner agency among at-risk cohorts." Miami Dade College with a "15% increase in student pass rates" and a 12 percent drop in course dropout rates. The University of South Florida with "55 to 60%" time savings on some tasks, with no statement of which tasks. The post says more than "17 million" people have completed an in-demand AI skills credential, a completion count rather than an outcome, and that AI-related skills command wage premiums "up to 40%," attributed to the AI Economy Institute with no underlying study named. Learner agency is not a standardized measure with an accepted instrument, so a 275 percent change in it has no fixed meaning even in principle, and no baseline, comparison group or instrument is identified. The agreement Microsoft signed nine days earlier is a real document with real clauses, and this is not that. The tell is that the most dramatic figure attaches to the least defined thing. When one of these lands in your inbox, do the boring thing: ask which of the five numbers came with a method, and notice that asking is usually enough to end the conversation. Microsoft →
Nicholas A. Gage, Research Director at WestEd, posted two companion preprints to EdArXiv on September 19, 2026, covering four elementary schools in one Indiana district that deployed an AI phonics tutor. Of 838 students in kindergarten through Grade 2, 97 percent were exposed to it, and the median student used it 6.7 minutes per week. Comparing high-dose users, 15 minutes a week or more, with near-zero users, fewer than 2 minutes a week, within the same schools and with baseline equivalence achieved through propensity weighting, produced a medium effect on the end-of-year DIBELS composite, Hedges's g of 0.47. A continuous dose-response within classrooms using teacher fixed effects found each additional 10 minutes per week predicted 7.1 composite percentile points across 43 classroom clusters, and the association was significantly larger for students who began the year with weaker reading skills and for students with disabilities. Among the 520 students who began below benchmark, each additional 10 minutes per week was associated with 2.40 times the odds of ending the year at or above benchmark. The three numbers below are the shape of the problem.
This week is about the difference between access and evidence. Florida wrote a parent's consent into rule and told districts to have a non-AI alternative ready. A randomized trial found ChatGPT saved science teachers 25.3 minutes a week on planning and changed nothing measurable about the resources, and never looked at students at all. Harvard's faculty mostly believe they have seen AI writing and mostly cannot prove it. Applicants whose essays were AI-written wrote better essays and were admitted less often. Microsoft published a 275 percent gain in something it never defines. And in four Indiana schools, 97 percent of children were given a reading tutor that the median child used for under seven minutes a week. In every one of these, the number that traveled was the easy one, and the number that mattered was the one nobody collected. Ask for the second one.
Quality AI lesson planning that lifts instructional quality and cuts teacher workload at the same time. It closes the gap between your curriculum and student-ready instruction, and gives leaders visibility into how your instructional model is actually being implemented, without piling more onto teachers.
See Planning Partner →