← All projects

Case study

Rung

One rung at a time.

A public AI math tutor for grade 6 and 7 ratios that gives hints instead of answers. The AI writes the words. Code does the grading and decides what the student is allowed to see.

Role
Product, design and build
Focus
Grade 6 and 7 ratios
Status
Live

The problem

Getting the answer is easy now. Learning still isn't.

Any student with a chatbot can get the answer to a ratio problem in seconds. That finishes the homework, but it skips the part where the student works something out and it sticks.

A good tutor does the opposite. They figure out where you went wrong, give you just enough to take the next step, and hold back the answer. I wanted to see whether an AI tutor could be built to behave that way reliably, and to prove it with measurements instead of a demo that happens to look good.

Who it's for

Middle schoolers working on ratios

Rung is for students in grades 6 and 7 practicing ratios. It's public and there's no account to create, so a student can open it and start right away.

Keeping it to one topic was deliberate. Ratios have well-known mistakes that students make again and again, which makes it possible to check the tutor's behavior closely instead of hoping it works across all of math.

What's already out there

Other AI and computer tutors

Khanmigo

Socratic AI tutor

Khan Academy's tutor asks guiding questions instead of giving answers. A University of Toronto study of 18 Tennessee middle schools, reported by the Hechinger Report on September 21, 2026, found that students mostly stopped using it.

Eedi Tutor

Closest in spirit

Built on data about student misconceptions and never gives the answer. Rung shares this idea: find the specific mistake, then respond to it.

Carnegie Learning MATHia

The lineage Rung borrows from

MATHia comes from the Cognitive Tutor line of research, which tracks what a student has mastered skill by skill. Rung's skill map uses the same family of methods.

The solution

A four-rung hint ladder

When a student gets stuck, Rung offers help in steps. Each rung gives a little more than the one before, so the full method comes last, not first.

1

Nudge

A light prompt to get the student thinking again.

2

Hint

A pointer toward the idea the problem depends on.

3

Example

A similar problem worked through, so the student can copy the approach.

4

Walkthrough

A full step-by-step path through the method.

Behind the ladder, code grades every answer and matches wrong answers to five known ratio misconceptions, so the help speaks to the mistake the student actually made. A skill map tracks mastery for each skill using Bayesian Knowledge Tracing. Before any reply reaches the student, code checks it and blocks it if it gives the answer away.

Evals

How I measured it

I wrote a 32-case eval that checks the things that matter for a tutor: does it leak the answer, does it make math mistakes, how fast does it respond, and does it sound like a person or a template. I ran the full set twice on September 30, before and after one prompt revision, and scored both runs with the same code.

0
answer leaks across 28 checked replies
0
math errors
~1.2s
median response time
32
test cases per run

What one prompt revision changed

MeasureBeforeAfter
Replies that opened with a stock phrase15 of 320 of 32
Walkthroughs that were complete3 of 44 of 4

The grader had a bug too

Partway through, I found a mistake in the grading code itself. It was counting the student's own answer as the tutor doing math. I fixed it, then graded the first run again with the fixed code so the before and after comparison is fair.

It was a useful reminder that evals are software too, and they need checking like anything else.

Key decisions

What I chose, and why

  • Code grades. The AI only writes the words.

    Whether an answer is right is a math question with one answer, so code decides it. The language model's job is to explain and encourage, which is what it's good at.

  • Hints in steps, not one big answer

    A fixed ladder of Nudge, Hint, Example and Walkthrough keeps the help predictable and saves the full method for last, after the student has had a chance to get there on their own.

  • Name the mistake

    Matching wrong answers to five known ratio misconceptions means the next hint is about the student's actual error, not a generic retry message.

  • Check every reply before it's shown

    A tutor that gives away the answer has failed, however friendly it sounds. Code checks each reply and blocks any that reveal the answer.

  • Track skills, not just scores

    A skill map built on Bayesian Knowledge Tracing estimates mastery for each skill, borrowing from the Cognitive Tutor research behind MATHia.

Guardrails

Safe to leave running in public

Rung is open to anyone, so it needed protection against misuse and runaway costs, and it needed to respect students' privacy from day one.

  • The API key lives only in Vercel, never in the browser
  • Per-visitor and daily limits enforced in code
  • A Vercel Firewall rule in front of the app
  • A monthly spend limit
  • No accounts, and nothing stored on the server

What I learned

Three things I'd carry into the next AI product

  • Decide what has to be right, and give that to code. Grading and answer checks don't belong to the model. With those in code, the AI can focus on what it does well: explaining.
  • Evals need checking too. The bug in my own grader would have made the tutor look worse than it was. Re-grading the first run with the fixed code is what made the comparison trustworthy.
  • Small prompt changes are only visible if you measure them. One revision took stock openers from 15 of 32 replies to 0. I wouldn't have known that from reading a few replies by hand.

Try it

See it for yourself

Rung is live. Try a few problems, get one wrong on purpose, and watch how it responds.