How to Use AI for Lesson Planning
Teachers using AI weekly report saving 5.9 hours a week. Where those hours really are — differentiation, rubrics, item generation — why you should start from objectives rather than activities, and what a model can never know about your class.
Teachers who use AI at least weekly save an average of 5.9 hours a week — about six weeks over a school year. That figure comes from a Gallup and Walton Family Foundation survey of 2,232 US public K-12 teachers, drawn from the RAND American Teacher Panel and fielded in spring 2025.
The same survey found only 32% of teachers use AI weekly. Four in ten don't use it at all.
Six weeks is a serious number, and it is worth being precise about where it comes from — because the hours are not saved uniformly across the job. They come from the documentation layer: the rubric, the worksheet variants, the differentiated version, the parent email, the unit map that exists mostly so someone can check it exists. That work is real and it is genuinely mechanical. It is also the part of teaching that expands to fill whatever time you give it.
Start from the objective, not the activity
The most common way AI lesson planning goes wrong is asking for the lesson.
Write me a lesson plan on fractions
produces something that looks complete and teaches nothing in particular. It will have a hook, three activities, and a plenary, and no coherent claim about what a student should be able to do at the end that they couldn't at the start.
Work in the other order:
- Objective first.
Turn this topic into 3-4 measurable learning objectives in student-outcome language, aligned to these standards: [paste].
Measurable means observable — identify, construct, justify — not understand or appreciate. - Assessment second.
For each objective, what specific student work would show it was met? What would a common near-miss look like?
- Activities last.
Now build activities that produce that evidence in a 50-minute period, with a hook and two checks for understanding.
This is backward design, and models are unusually good at it because each step gives the next one hard constraints to satisfy. Ask for activities first and the model has nothing to aim at, so it produces plausible-looking classroom filler.
The teaching pack is sequenced this way on purpose — objectives, then rubrics and assessments, then lesson plans, units, and materials.
Where the hours actually are
Ranked by time saved per unit of judgement required:
Differentiation. One worksheet at three levels, or one text at three reading levels with the key vocabulary held constant. This is the strongest case in the whole list: it is pure transformation, you can verify it at a glance, and it is the thing most likely to get skipped at 10pm.
Rubrics. Criteria, performance levels, and descriptors specific enough that two markers agree. Ask for the descriptors to be behavioural — what the work does, not how good it is — and the rubric becomes usable by students before they submit, which is where rubrics earn their keep.
Item generation. Twenty questions at a stated difficulty with an answer key. Always check the key. Models produce confident, wrong answer keys at a rate that will surprise you, particularly on multi-step arithmetic and on questions where a distractor is defensible.
Explanations at multiple levels. The same concept as a one-sentence definition, a plain-language walkthrough, and an analogy — so you have something to reach for when the first explanation doesn't land.
What it cannot do, stated plainly
It doesn't know your class. Every plan it produces is for a generic room. Which three students need the scaffold, who will finish in four minutes and disrupt, what happened last Tuesday that means you cannot use that example — none of it is in the model and none of it can be. The plan is a starting draft that a person who was in the room has to finish.
It doesn't know your standards. It will produce standard codes that look right and are wrong. Paste the actual text of the standard you are aligning to. Do not ask a model to recall a curriculum framework from memory.
It grades on surface features. Against a rubric it will reliably reward fluent, well-organised, confidently-worded writing. A muddled paragraph containing a genuinely original insight is exactly the case it handles worst — and exactly the case that most deserves a teacher's attention. Use it to triage and to draft comments; do not let it be the mark.
Designing what students use, not just what you use
There is a second finding every teacher deploying AI to students should know. In a PNAS field experiment, Bastani and colleagues gave nearly a thousand high school maths students either a standard GPT-4 chat interface or one prompted to give hints rather than answers. Both improved practice performance dramatically. But when the tools were removed for the exam, the standard-interface students scored 17% worse than students who never had AI at all. The hint-based version largely prevented that.
The design choice — hints versus answers — was the whole difference between a tool that helped and a tool that hurt. If you are putting an AI tutor in front of students, that single instruction matters more than which model you pick.
A realistic view of the six weeks
The Gallup figure is self-reported time savings from teachers who already chose to use AI weekly, which is not the same as a measured effect on a random teacher. Take it as a real signal about where the tedium lives, not a guarantee.
The honest version is narrower and still worth having: the parts of the job that are documentation get faster, and the parts that are teaching do not. What you do with the recovered time is the actual question — and it is a better question than the one about which tool to use.
Sources: Gallup and Walton Family Foundation, Teaching for Tomorrow: Unlocking Six Weeks a Year With AI
(June 2025), n=2,232 US public K-12 teachers via the RAND American Teacher Panel, fielded 18 March – 11 April 2025. Bastani et al., Generative AI Without Guardrails Can Harm Learning,
PNAS (2025).