ARTICLES BIAIS ET ETHIQUES IA

Generative AI in Teaching: What to Delegate and Where Human Judgement Matters

Generative AI can help teachers prepare, interact with learners, give feedback and even evaluate. This practical five-level framework helps teachers and trainers decide what to delegate, where human judgement matters and what European AI rules already require.
Teacher using generative AI while keeping human judgement and oversight in the learning process

Generative AI in Teaching: What to Delegate and Where Human Judgement Matters

In one sentence: AI can help you prepare a session, or it can judge a learner's work. Those are very different uses. The closer AI gets to the learner, the clearer you need to be about who is still responsible.

Key takeaways

  • Preparing an exercise and marking a learner are not the same thing. The closer the system gets to the learner, the more it can affect their results.
  • This article offers five levels of delegation. We built them to help you see how far you have handed a task over.
  • European rules already apply to education. Schools and training providers must help their staff understand the AI tools they use. Some educational uses count as high-risk. Using AI to work out how learners feel from biometric data — their face, voice or body — is banned in schools and training settings, apart from narrow medical and safety cases.
  • Keep asking one question. What can this output change for the learner, and who checks it before it reaches them?

What problem does generative AI create in teaching?

It usually starts on a Sunday evening. You have a session tomorrow, you are short of examples, and you ask an AI for three versions of an exercise. You read them, you keep one, you change two lines. Nothing about that feels like a decision worth writing down.

A few months later, the same tool is answering learners' questions between sessions. Then it is commenting on their drafts. Then someone suggests it could do a first pass on marking.

All of this gets called "using AI in teaching", and the phrase covers acts that have almost nothing in common. You can check a handout before anyone reads it. You cannot check a chatbot conversation that happened at eleven at night. And when a system writes "this argument is weak", the learner has no way of knowing who decided that, so the comment reads like a judgement rather than a suggestion. A recommendation then starts to decide what a learner does next.

Before we address the rules, there is a question we need to answer: who checks the output before it reaches the learner?

What is delegation in an AI teaching context?

Delegation, here, means handing an AI system part of a job you would otherwise do yourself. The closer it gets to your learners, the more you need to spell out what it does, where it stops, and who is watching it.

Here are two examples. You ask a model for three case studies, read them, rewrite one, and use it in the room. The system helped you, and you stayed between it and your learners. Now you add an assistant that answers their questions between sessions. The system is doing part of the teaching itself, and you are not in the room when it happens.

Neither one is good or bad in itself. They just need different safeguards. The confusion starts when you treat them as the same thing.

Harvard Business School has gone further than most. Its own page describes Foundry's Bootcamp as an eight-week programme for founders, powered by an AI platform. What the press added in August 2026 is the part that makes people stop: the course is priced at $699, and the feedback during practice pitches comes from AI avatars of the faculty, built with the video company HeyGen. Harvard told Fortune that around 760 founders had been through it. Most of us will never face that decision. We face the small ones, over and over. We examine that case in Harvard's AI Faculty Avatars: Who Controls the Digital Double?

What is the difference between AI assistance and AI delegation?

"Assistance" and "delegation" are two words for something that happens in stages. You hand a task over bit by bit. Here are five steps. The important question to ask is which one you are on today.

Level 1. AI helps you

Outlines, examples, exercises, summaries. Your learners never meet the system.

What you gain: time, and more variety than you would produce on your own.
What you risk: confident mistakes, and material you did not really choose ending up in your session.
What stays with you: reading it properly before anyone else does.

Level 2. AI talks to your learners

A chatbot, a revision assistant, a tutor answering directly.

What you gain: availability you cannot offer, and a place where someone can ask the question they were too embarrassed to ask you.
What you risk: a wrong answer delivered confidently, with nobody there to catch it.
What stays with you: saying what the assistant is for, what it is not for, and when to come and find a human.

European law backs this up. Since August 2026, a system built to talk directly to people has to make clear that it is AI, unless that is already obvious from the situation. The European Commission's guidance names chatbots, AI agents and avatars among them.

Level 3. AI comments on their work

It reacts to a draft, an exercise, a recorded presentation. This is where things get genuinely blurry.

What you gain: learners get a reaction straight away and can improve before showing you anything.
What you risk: a learner reads "this argument is weak" and takes it as the final word, even though no teacher has checked whether the comment is right.
What stays with you: deciding whether this is help or assessment, and telling your learners which.

Level 4. AI marks, or decides what comes next

It scores. It feeds progression decisions. It shapes which module someone is offered.

What you gain: a steadier process when there are a lot of learners, and a record of how each was assessed.
What you risk: the system's answer now affects a real person, and nobody has checked how reliable that answer is.
What stays with you: putting your name on the decision, and being there when a learner wants to argue with it.

Level 5. AI speaks as you

An avatar, a synthetic voice, a digital copy answering in your name.

What you gain: a presence that does not depend on you being there.
What you risk: your name and your credibility attached to answers you never gave and may never see.
What stays with you: agreeing what it may say, who controls it and how long it lasts — before anything is recorded.

One thing to be clear about

We built these five levels ourselves. You will not find them in any regulation. A level does not tell you whether a system is lawful, high-risk or a good idea. It tells you how far you have handed a task over. You need to know that before you can answer any of the rest.

Why does this become a governance question?

Governance sounds like a word for big organisations with committees. In practice, it is five simple questions. What is this system allowed to do? Who can it talk to? Can it only draft, or can it also mark and record? Who reads what it produces? And who can switch it off?

European rules already answer part of that for you.

Regulatory note

AI literacy. Since February 2025, two groups have had a duty here: the companies that supply AI systems, and the organisations that use them. Both must take steps to build AI literacy among their own staff and among anyone who operates the system for them. They should take account of what those staff already know and of how the system is used. The text was rewritten in July 2026 and added one clarification: you do not have to prove that each person has reached a given level of knowledge. You have to make the effort, not certify the result. Supervision and enforcement started in early August 2026. (AI Act, Article 4, as amended by Regulation (EU) 2026/1744.)

High-risk uses. The AI Act lists four education and training uses as high-risk. A system is on the list if it decides who gets admitted, judges what a learner has achieved and that result shapes their path, decides what level of education suits someone, or watches for cheating during tests.

Being on the list does not settle the question. A system can still fall outside the high-risk category if it only performs a small routine step, or prepares work that a person then decides on, and does not really change the outcome. One exception holds in every case: if the system profiles individuals, it stays high-risk. These requirements apply from 2 December 2027. (AI Act, Annex III point 3 and Article 6(3); timing set by Regulation (EU) 2026/1744.)

Human supervision. If your organisation uses a high-risk system, the law does not ask you to have a human somewhere nearby. It asks you to assign supervision to named people who have the competence, the training and the authority to do it, plus the support they need. (AI Act, Article 26(2).)

Using AI to work out how someone is feeling, from their face, voice or body, is banned in schools, training settings and workplaces. The only exceptions are narrow medical and safety ones. This has applied since February 2025.

The reason is worth knowing, because it is not just about consent. The AI Act states that the scientific basis for reading emotions is weak. People show emotion differently from one culture to another, from one situation to another, and even from one day to the next. The Act calls these systems unreliable, and hard to trust outside the cases they were built for, which is how you end up with unfair results. The law then adds what we already know: a learner or an employee is not in a position to say no.

Elsewhere

Other countries frame the same problem differently, and this is our reading of it rather than an official classification. In the United States, the argument is about digital copies of real people: what the person agreed to, and who controls the copy. A federal bill has moved through the Senate without becoming law, so protection still depends on state rules that differ from one state to the next. In China, rules in force since September 2025 require AI-generated content to carry both a label you can see and technical information inside the file, so it can still be recognised further down the line.

Where does the problem show up in practice?

Imagine a training provider that adds an AI assistant so learners can ask questions between sessions. Six months later, trainers are pasting learners' assignments into the same assistant to get a first round of comments. Marking takes hours, and the comments are good.

Nobody changed anything. No new tool, no new contract, no new policy. But the system has moved from level 2 to level 3. We suddenly notice that comments start feeding a real assessment. The quality of the comments is not the issue. The issue is that a new use was never agreed, never mentioned to learners, and never reviewed, because every single step looked sensible on its own.

The same thing happens in other ways. Feedback turns into a mark when a comment written to help someone improve is pasted into a progress report. Suggestions turn into gatekeeping when a system that recommends modules starts deciding who gets offered what. Monitoring turns into mind-reading when a tool stops tracking activity and starts inferring feelings, which is where the ban applies. A recording outgrows its permission when material captured for one course is turned into an interactive avatar for another.

In each case, the tool has not changed. Its role has. This is what we call purpose drift.

Why does human judgement matter?

"There is a human in the loop" is meant to be reassuring. On its own it tells you almost nothing. It becomes useful when you identify which person, for what task, and with what power.

Four moments are worth naming. Before you start, decide what the system may do and what it must never decide on its own. While testing it, try the difficult cases rather than the clean ones: hesitant writing, answers in a second language, replies that are partly wrong but show real understanding. While it runs, make sure someone can see what it is really producing, not a sample chosen for them. When reviewing, give that person the power to change it, limit it or stop it.

That last one is where supervision usually turns out to be thinner than it looked. Someone who can report a problem but cannot fix it is not supervising anything. European law says the same thing about high-risk systems, in its own words: the people responsible need skills, training, authority and support.

Quick check

Who is allowed to stop or restrict this AI use — and would that person know it is part of their job?

Human judgement matters most when an output can change a learner's mark, their progression, what they get offered, or how they see their own work. Get that clear and the rest gets easier: preparing material, generating alternatives, running a role-play, explaining something a third way.

How can a teaching team check this in 15 minutes?

Delegation check for teachers and trainers

Pick one specific use, not "AI" in general. Count only the answers you can state clearly to a colleague.

  1. Which level are we at?
    Helping us prepare, talking to learners, commenting on their work, marking, or speaking in someone's name. Name the action precisely: "it helps with marking" covers too much.
  2. What are we putting into it?
    Anything personal or confidential about a learner deserves its own decision before it goes anywhere near the system.
  3. What do learners think it is doing?
    Clear to them while they are using it, not buried in a document nobody opens.
  4. Can the output affect a mark, an opportunity or a path?
    Including indirectly, by shaping what a human decides.
  5. Who checks it, and who can stop it?
    A name and a role, with the authority to act rather than only to flag.
  6. Do the people using it know enough?
    Including the colleagues who inherited the tool instead of choosing it.
  7. What happens when the use changes?
    A new purpose should mean going through these questions again. This is where purpose drift usually happens.

One limit: this is a conversation starter, not a compliance check. It does not replace legal advice, a data protection assessment, a technical review or your sector's own rules. A low score does not mean your tool is unsafe. It means decisions are still being made by default.

Conclusion

Most arguments about AI in teaching happen because the two people are picturing different things. One is thinking about help with a lesson plan. The other is thinking about a machine marking their students. Both say "AI".

So before bringing in a new tool, or letting an existing one do more, ask three questions. Which level are we at? What can the output change for a learner? And who has the authority to step in?

The answer does not have to be less AI. It has to be a clearer decision.

Bringing this to your team?

Prompt & Pulse helps organisations adopt and use generative AI with clearer rules, responsibilities and safeguards: diagnosis, AI acculturation across teams, and hands-on support, with a particular focus on bias, ethics and responsible use.

For teachers, trainers and learning teams, that means working through what gets delegated to AI, who supervises it, who remains responsible, and how AI outputs reach learners. Available in French and English.

Start a conversation

FAQ

Does the AI Act require teachers to be trained in AI?

Not in those words. The duty sits with the organisation that supplies or uses the system: it has to take steps to build AI literacy among its staff and anyone operating the system on its behalf. Since the July 2026 amendment, the text says plainly that no particular level has to be guaranteed for any individual. It is an obligation to make the effort, not a qualification your teachers have to pass.

Is an AI feedback tool automatically high-risk?

No. Systems meant to judge what a learner has achieved can fall into the high-risk category, but it depends on what the tool was built for and what its output actually does in your process. A system that only handles a narrow procedural step, without really shaping the outcome, can fall outside it — unless it profiles individuals, in which case it stays high-risk. Your specific tool needs looking at, not a rule of thumb.

Can we use AI to check whether learners are paying attention?

It depends what it measures. If it works out how someone feels from their face, voice or body, that is banned in education, with narrow medical and safety exceptions. The reasoning goes beyond privacy: the law points to weak scientific grounds, unreliable results, the risk of unfair outcomes, and the fact that a learner is in no position to say no. Other ways of measuring engagement need their own look at the law, the ethics and the teaching value.

Do learners need to be told when they are dealing with AI?

For systems built to talk directly to people, yes, unless it is obvious from the situation — and the Commission's guidance names chatbots, agents and avatars. Beyond the rule, there is a plainer reason: someone who does not know where an answer came from cannot judge how much to trust it.

What is the difference between assistance and delegation?

With assistance, you stay between the system and the learner. With delegation, the system does part of the job itself: talking to them, commenting on their work, shaping their assessment, or standing in for you. The five levels are just a way of saying where a given use really sits.

Should teachers avoid generative AI?

Nothing here supports that. Preparing, explaining, practising, commenting, marking and standing in for someone carry very different consequences, and they deserve to be judged separately rather than lumped together.

Is this only about schools and universities?

No. Corporate training, certification, coaching and onboarding raise the same questions, often with less supervision than a school would have. What the law requires depends on your context and on the system.

Sources and references

  1. Regulation (EU) 2024/1689 (AI Act): Article 4 (AI literacy), Article 5 (prohibited practices, including inferring emotions in workplaces and education institutions), Article 6(3), Article 26(2) (human oversight assigned by deployers), Article 50 (transparency), Annex III point 3 and Recital 44. Full text on EUR-Lex.
  2. Regulation (EU) 2026/1744 of 8 July 2026, Digital Omnibus on AI, in force since 27 July 2026: amends Article 4 and sets 2 December 2027 for the high-risk requirements applying to Annex III systems. Official Journal text.
  3. European Commission, AI Act Service Desk, Article 26, deployer obligations, and AI Literacy: Questions and Answers, updated July 2026.
  4. European Commission, Guidelines on transparency obligations for providers and deployers of AI systems under Article 50, final version, 20 July 2026.
  5. Harvard Business School, Foundry Bootcamp page (eight-week programme, AI platform).
  6. Anthony Ha, "Harvard's $699 startup bootcamp offers AI avatars of its instructors", TechCrunch, 22 August 2026, and Fortune, August 2026, for the price, the avatars and the number of participants.
  7. Paul Baier, "6 Lessons From The Harvard Business School AI Project For Founders", Forbes, 24 July 2026, on how the Foundry product was built and tested.
  8. Cyberspace Administration of China, Ministry of Industry and Information Technology, Ministry of Public Security and National Radio and Television Administration, Measures for Labelling AI-Generated Synthetic Content (notice 国信办通字〔2025〕2号, issued 7 March 2025, published 14 March 2025): explicit and implicit labelling, Articles 3 and 5; duties of distribution platforms, Article 6; in force 1 September 2025, Article 14. Unofficial English translation: China Law Translate.
  9. U.S. Copyright Office, Copyright and Artificial Intelligence, Part 1: Digital Replicas, 31 July 2024, recommending a federal law against unauthorised digital replicas.

Transparency note: This article was co-written with the assistance of generative AI Claude AI and ChatGPT. The illustration accompanying it was also generated with AI. The author defined the angle, selected and reviewed the sources, challenged the analysis and validated the final version. AI-generated suggestions were not treated as evidence. Legal claims were checked against the official EU sources listed above, and the Harvard facts are separated between what the school publishes itself and what the press has reported. The five-level framework and the seven-point check are editorial tools developed by Prompt & Pulse, not legal classifications. Where applying a rule depends on facts an article cannot know — in particular whether a given feedback tool counts as judging what a learner has achieved — the article says so rather than guessing. European legislation in this area is moving quickly, including on dates that have already shifted once, so every date should be checked again before publication. This article does not constitute legal advice. The author specializes in AI ethics, bias detection and responsible deployment for SMEs.