Act One: The Record Button
Mei stared at the red record button on her screen, finger hovering in mid-air.
The classmate beside her had already pressed it. So had the one behind her. English sentences filled the room from every direction, all sounding more fluent than anything she could produce. Mei angled her phone downward, closer to her mouth — but really, she was trying to make sure nobody else could hear.
She rehearsed the sentence three times in her head. The first attempt stalled on the second word. The second had completely wrong intonation. The third — forget it. The teacher would hear it anyway. Better not to record at all.
She pressed cancel.
This wasn't the first time. Last week she'd summoned the courage to submit a recording. The next day, the teacher's feedback arrived: "Pronunciation needs work. Intonation sounds unnatural." The feedback was accurate. But Mei didn't practice again for three days.
Not because she was lazy. Because she was ashamed.
Act Two: The Teacher's Office, 9 PM
On the other side of the screen, Ms. Chen rubbed her eyes. Twelve recordings left to grade.
She'd already done twenty-eight today. Each one took three to five minutes — listen to the student speak, run multiple evaluation threads simultaneously in her head: Is the pronunciation correct? Does the intonation sound natural? Is the speech fluent? Did they actually answer the prompt? Then compress all those judgments into feedback a student can understand.
Forty recordings a day. Every single one demanding full attention.
She wasn't slacking. She cared too much. That was exactly why she felt like she was turning into a grading machine.
"Recording three — too many ums, watch the intonation." "Recording four — off topic, points deducted." "Recording five — same issue as recording three."
Ms. Chen remembered why she'd chosen this job in the first place: because watching a student open their mouth and speak English for the first time — that moment was genuinely beautiful.
But she hadn't felt that moment in a long time.
Act Three: A Phone Call
A friend messaged me one day: "I know someone running an English speaking platform. Students are growing but teachers can't keep up. Think you can help?"
The call lasted thirty minutes. The founder was direct: students record themselves speaking English, teachers grade and give feedback. Good word-of-mouth. But each teacher was handling forty recordings a day. Some students waited two or three days for feedback — by which point the motivation had already evaporated.
He wanted AI to handle the first pass. Teachers would only review what the AI flagged as uncertain.
I agreed. But I had one condition.
"Before I write a single line of code, let me sit with your teachers and watch them grade."
Act Four: An Afternoon of Observation
I spent an entire afternoon sitting beside two teachers.
I asked them to narrate their thought process aloud as they listened to each recording. This wasn't an interview. It was more like transcribing someone's mental operating system, line by line.
When a teacher listens to a recording, several things are running in parallel:
- Pronunciation: Did the student say the word correctly? Are the vowels off?
- Intonation: Does it sound like reading from a textbook, or like actually speaking?
- Fluency: Are they constantly stumbling? Too many ums and uhs?
- Topic relevance: Did they answer the question? Or did they talk around it without a point?
And the weighting isn't fixed. For beginner students, pronunciation matters most. For advanced students, content and fluency take priority.
None of this was documented anywhere. It lived in the teachers' intuition, compressed by thousands of hours of teaching into instinct.
My job was to extract it and translate it into rules a machine could execute.
Act Five: The Sixty-Recording Blind Test
After the system was built, I designed an exam — not for students, but for the AI.
Sixty student recordings. Teachers graded all sixty. The AI graded the same sixty. Both sets of results laid side by side, compared line by line.
If this were a school drama, this is the final exam scene. The classroom is silent. Everyone is waiting for the report card.
The report card came back. The AI did not pass.
Too lenient on fluency. Students who stuttered constantly, who packed their speech with "um" and "uh" — teachers docked points immediately. The AI thought "the words are all correct" and barely noticed. It heard every syllable but couldn't hear rhythm or confidence.
Too literal on relevance. A student didn't answer the prompt word-for-word, but the overall meaning was there. Teachers gave credit for intent. The AI saw only deviation.
Too polite in feedback. The AI wrote like a greeting card: "Overall you did well, but consider improving X." Teachers just said: "This is wrong. Fix it like this."
These weren't technical problems. They were problems of educational judgment. The AI had learned to listen, but hadn't learned when to be strict and when to let go.
Act Six: The Retake
Over the next several rounds, I fixed the gaps one by one.
Each adjustment was like teaching the AI a new concept. "When a student pauses in this context, how does a teacher respond?" "When a student goes off-topic but the meaning lands, what's a fair score?" I wrote in dozens of specific cases — not rules, but scripts from the teaching floor.
Ran the blind test again. The gap narrowed. Again. Narrower still.
Until one day, a teacher read the AI's feedback and said:
"Yeah. That sounds like something I would write."
When the founder relayed this to me, he added: "That was the moment I knew you got it right."
The AI passed its retake. Not because it was perfect — but because its imperfections had become close enough to a teacher's imperfections that the difference no longer mattered.
Act Seven: The Sentence That Stayed with Me
A few weeks after launch, the founder sent me one number.
Students were practicing speaking three times more often.
He'd asked a few students why. One answer stuck with him — and with me for a long time:
"Because I'm not afraid of the teacher hearing me anymore."
Students like Mei weren't unwilling to practice. They were afraid. Afraid of sounding terrible in front of a real person — a person who would remember how bad they were last time.
The AI doesn't laugh at you. Doesn't remember yesterday's mistakes. Doesn't give you that look — the one that says haven't we been over this?
So they could try, fail, and try again. In front of the AI, mistakes carried no social cost.
I never wrote "remove fear" as a feature in any spec. But it turned out to be the most important output of the entire system.
Act Eight: The Teachers Didn't Disappear
Teacher grading volume dropped from forty recordings a day to about fifteen. But they weren't working less.
They became the AI's coaches. When the AI produced a strange evaluation, teachers flagged it. The next time a similar case arose, the AI did better. Teachers started spending their time where the AI genuinely couldn't help — students who said something truly bizarre, students who were clearly struggling emotionally, students who needed encouragement rather than correction.
Ms. Chen told me something I haven't forgotten:
"Before, I graded forty recordings a day. Now it's about fifteen. But I feel like I actually know my students better."
She was feeling that moment again — the one where a student opens their mouth to speak English and crosses from fear to not-fear. Only now, she had time to actually see it.
Epilogue: The Real Value of AI in Education
One thing this project taught me:
The biggest value AI brings to education isn't replacing teachers. It isn't cutting costs.
It's making students unafraid to fail.
The hardest part of learning has never been the knowledge itself. It's "am I willing to expose what I don't know?" AI created a space where you can fail endlessly, with no one judging you.
Traditional classrooms can't do this. Not because teachers don't want to — but because the pressure between people is inherent in any human relationship.
And a system without emotions, precisely because of that absence, became the safest practice partner of all.
Mei presses that red record button every day now. Sometimes three times a day.
She doesn't need to summon courage anymore.
I'm Young. I help organizations build their digital stack — from process mapping to AI integration, end to end. If you're working on something similar, I'd love to chat.
