Independent portfolio project
AI Learner Diagnostic
An independent prototype for diagnosing a learner’s skill gaps and proposing the next step, designed around what happens when the AI is unsure or wrong.
- Outcome
- Independent prototype · deterministic demo, no real users or model results
- Focus
- AI productEvaluationHuman oversight
- Evidence
- Independent prototype
Independent portfolio project, not shipped at any employer. The interactive demo runs on deterministic logic so the experience is reliable and transparent. No real-user adoption, model accuracy or production results are claimed.
The 30-second version
Read the full story1Problem
Working out what a learner is missing and what they should do next is judgment-heavy work that teachers rarely have time to do for every student.
2Approach
Split the job between rules, a model and the teacher; designed confidence, fallbacks and override into the UX; defined the evaluation and launch gate before any model work.
3Outcome
A working, deterministic prototype of the product loop. No real-user adoption or model-performance results are claimed.
Why AI
Framing
Problem first, model second.
Working out what a learner is missing, and what they should do next, is judgment-heavy work. Teachers do it well and rarely have time to do it for every student. The gap is not content. It is diagnosis at scale.
It is also a problem where being wrong is costly in a quiet way. A learner sent down the wrong path doesn’t complain, they just stop. So the design question was never whether a model can do this. It was where a model should do it, and what happens when it is wrong.
The diagnostic loop
Product reasoningAssess
Collect a small, representative set of learner evidence.
Diagnose
Identify the skill gaps the evidence actually supports.
Explain
Show the learner and educator why, with evidence before the verdict.
Recommend
Propose a learning path the educator can accept or override.
Practice
Generate targeted practice for the diagnosed gap.
Evaluate
Score the outcome against a rubric, not a demo prompt.
Adapt
Feed the result back into the next decision.
- ↺ back to step 1
A loop, not a one-shot answer. Each step is a product surface with its own failure mode, which is why explanation and evaluation are steps in the loop rather than afterthoughts.
Prototype
Interactive prototype
Run it. Then try to break its confidence.
Choose answers that ignore the learner signal and watch the output change. Confidence drops, the evidence trace shows why, and at low confidence the system stops guessing and hands the decision to the educator.
Learner Diagnostic
Demo mode · deterministic
Try it · 3 signals, ~1 minute
See the diagnostic loop, not just the architecture.
Answer three representative learner-signal questions. The prototype turns those signals into a transparent diagnosis, states how confident it is, and hands the final call to the educator.
This is a deterministic product prototype, not a claim of production model performance. The production architecture below would use an LLM, retrieval, evaluation data, and human controls.
Rules vs. model
The first decision
Where rules win, and where a model earns its place.
The first design decision was a split, not a model choice. Anything that has to be consistent, auditable or cheap stays deterministic. The model gets the parts that need judgment over messy evidence, and each of those parts has a defined fallback. The teacher keeps the final call.
Proposed production split
Product reasoning| Job | Owner | Why |
|---|---|---|
| Score answers | Rules | Deterministic and testable. There is nothing to guess. |
| Combine accuracy, hints and time into a readiness signal | Rules | The same input must give the same output, and a teacher must be able to check it. |
| Map a pattern of errors to a likely misconception | Model | Needs judgment across messy evidence. Rule-based in the prototype. |
| Explain the diagnosis to learner and teacher | Model | Language generation, grounded in the evidence trace. |
| Generate targeted practice | Model + bank | Variety, anchored to a curated, vetted practice bank. |
| Decide what happens at low confidence | Rules | A fallback has to be predictable. |
| Accept or change the plan | Teacher | Accountability stays with a person. Overrides are logged. |
Score answers
RulesDeterministic and testable. There is nothing to guess.
Combine accuracy, hints and time into a readiness signal
RulesThe same input must give the same output, and a teacher must be able to check it.
Map a pattern of errors to a likely misconception
ModelNeeds judgment across messy evidence. Rule-based in the prototype.
Explain the diagnosis to learner and teacher
ModelLanguage generation, grounded in the evidence trace.
Generate targeted practice
Model + bankVariety, anchored to a curated, vetted practice bank.
Decide what happens at low confidence
RulesA fallback has to be predictable.
Accept or change the plan
TeacherAccountability stays with a person. Overrides are logged.
In the prototype every row runs on deterministic rules, so the demo is reliable and inspectable. This table is the proposed production split.
System design
System thinking
Where the model sits, and where it doesn’t.
A practical architecture combines structured learner signals, retrieval from a curated knowledge base, an LLM for the judgment-heavy steps, deterministic scoring wherever possible, and a feedback and evaluation loop.
01 · Inputs
- Structured learner signals
- Answers, attempts, hints, time on task
02 · Grounding
- Retrieval from a curated knowledge base
- Skill map and practice bank
03 · Reasoning
- LLM proposes diagnosis and next step
- Deterministic scoring where possible
04 · Experience
- Evidence, confidence and explanation
- Educator accept or override
05 · Learning loop
- Evaluation dataset and rubric
- Overrides logged as feedback
Failure modes
Failure modes
What happens when it’s wrong.
Every AI feature fails. The product decision is how: what the user sees, how the failure is detected, and what the system does instead. Designing these before the happy path keeps the demo honest.
Failure modes, designed before the happy path
Product reasoning01
Overconfident diagnosis
Looks like: A firm verdict from two answers
Caught by
Too little evidence for the confidence claimed
Product does instead
State low confidence, hold the current path, ask for more evidence
02
Hallucinated gap
Looks like: A skill the learner was never tested on
Caught by
Every claim must trace to an answer in the evidence
Product does instead
Drop untraceable claims before anything is shown
03
Contradictory signals
Looks like: Fast, correct answers with heavy hint use
Caught by
Signals disagree beyond a set margin
Product does instead
Flag it for the teacher instead of picking one reading
04
Discouraging language
Looks like: “You are weak at fractions”
Caught by
Tone checks in the evaluation set
Product does instead
Describe the gap and the next step, never the learner
05
Slow or failed model call
Looks like: A learner waiting at a checkpoint
Caught by
Latency budget exceeded or timeout
Product does instead
Serve the rule-based next step and diagnose in the background
06
Teacher disagrees
Looks like: An override
Caught by
Every override is logged
Product does instead
The override wins, and becomes a new evaluation case
Evaluation
Evaluation
Evaluation is a product practice.
Before launch, evaluate diagnostic accuracy, recommendation relevance, groundedness, harmful or overconfident outputs, consistency, latency and cost, against a representative evaluation set rather than a few demo prompts.
Pre-launch evaluation rubric
Product reasoning| Criterion | The question we test | If it fails |
|---|---|---|
| Diagnostic accuracy | Does the diagnosis match what an expert educator would conclude from the same evidence? | Wrong path for the learner |
| Recommendation relevance | Is the next step the most useful one for this gap, at this level? | Busywork and disengagement |
| Groundedness | Is every claim traceable to learner evidence or curated content? | Hallucinated gaps |
| Overconfidence and harm | Does it hedge when evidence is thin, and avoid discouraging language? | Loss of learner and educator trust |
| Consistency | Do similar learners get similar diagnoses across runs? | Unpredictable experience |
| Latency and cost | Is it fast and cheap enough to run at every checkpoint? | Unviable unit economics |
Diagnostic accuracy
Does the diagnosis match what an expert educator would conclude from the same evidence?
If it fails: Wrong path for the learner
Recommendation relevance
Is the next step the most useful one for this gap, at this level?
If it fails: Busywork and disengagement
Groundedness
Is every claim traceable to learner evidence or curated content?
If it fails: Hallucinated gaps
Overconfidence and harm
Does it hedge when evidence is thin, and avoid discouraging language?
If it fails: Loss of learner and educator trust
Consistency
Do similar learners get similar diagnoses across runs?
If it fails: Unpredictable experience
Latency and cost
Is it fast and cheap enough to run at every checkpoint?
If it fails: Unviable unit economics
Guardrails
Guardrails
Designing for uncertainty.
Show evidence where it exists, never present uncertainty as certainty, let a person override, and define what the system does when learner signals are thin or ambiguous.
Show the evidence
Every diagnosis lists the signals it used, so it can be checked.
Honest confidence
Uncertainty is stated, never styled away.
Human override
The educator makes the final call; overrides are logged as feedback.
Defined failure state
When signals are ambiguous, the system holds instead of guessing.
Launch criteria
Launch criteria
A gate, not a date.
Ship only when quality thresholds hold across representative cases and the AI experience measurably beats a credible non-AI baseline.
The first real test would be narrow: one subject, a few teachers, and the prototype’s diagnosis compared with the teacher’s own call on the same evidence. If it can’t agree with teachers often enough to save them time, it isn’t ready, however good the demo looks.
Launch gate: all four must hold
- Quality thresholds met across representative evaluation cases, not a handful of demo prompts
- No harmful or overconfident outputs in the failure-category review
- Latency and cost viable at every learner checkpoint
- Measurable improvement over a credible non-AI baseline
“Ship when it beats a credible non-AI baseline, not when the demo looks good.”
What this prototype is, and isn’t
It demonstrates the product loop and the decisions around it. The demo uses deterministic logic so it behaves the same way every time; a production version would connect the same flow to an LLM, curated retrieval, an evaluation set and human controls. It has no real users, and no adoption or model-performance results are claimed.