How Shello teaches
Learning that lasts is work you do yourself. You have to think the idea through, come back to it as it fades, and know honestly what you can't yet do. Most learning apps are built for the opposite: an answer now, a summary written for you, a score that flatters you. Shello is built around the three harder things. This page says why.
The learning path
Three things happen, in order
The tutor teaches from a real book, and asks before it tells
So a lesson opens with a question about a chapter you can read yourself, and the evidence for starting there is unusually strong. Researchers compare teaching methods on one scale, called effect size. On it, one-to-one human tutoring scores about d ≈ 0.79 — one of the largest effects known, and smaller than the figure usually quoted. Software that tutors step by step scores about the same as a human tutor.
In a 2025 randomized trial, students learned more, in less time, from an AI tutor built to enforce good teaching than in an active-learning class. The design did the work. That is Shello's bet too: the teaching lives in how the system is built, never in how clever the model is. Open chat with a model is not what the trial tested.
What you learn is then scheduled against forgetting
Understanding something once is not the same as still knowing it a month later. Two study techniques stand clear of the rest in the research: testing yourself, and spreading those tests out over time. Shello's cards do both, so each concept comes back as a question around the time you would otherwise lose it. The same research cuts the other way: spaced practice makes whatever you rehearse stick, right or wrong. That is why nothing goes on a card until it has been checked against the chapter it came from.
The one place we won't bend: honesty about what you know
The mastery map — the record of what you can and can't yet do — has to be honest even when that's unwelcome. Learners who overrate what they know measurably learn and retain less: they ignore feedback and skip the study that would have fixed it. So the map shows what your answers show, and a weak concept looks weak. Everything else — pace, mode, how far you want to go — is yours to set, and we never use guilt to steer you.
See it in practice
A lesson begins before the answer
Tutor waits for your reasoning, then uses the chapter to test and refine it. The question is not a delay before teaching; it is where the teaching begins.
A value far above the rest is added to a small data set. Before calculating, what do you expect it to do to the mean — and why?
Two myths from our own field that we won't repeat
The first is the figure usually quoted for tutoring. Bloom's famous "2-sigma" claim — that one-to-one tutoring lifts students by 2 standard deviations — is folklore in its strong form. The reviewed figure is d ≈ 0.79; the myth inflates it about 2.5 times. We build on the honest number.
The second is about how people study. Summarizing, highlighting, and rereading rate among the least useful techniques in the research, and they are the ones most learning apps automate. That is why Shello has no auto-summary button. The book gets you to pull the ideas together yourself; it never does it for you.
The evidence, decision by decision
Everything above is the short version. Below is the whole audit: each decision, what we bet on it, and the verdict. "Supported" means published studies back the choice; "practice standard" means professional norms do, where no controlled study exists or applies. The audit is a document kept with our code. The table is rebuilt from it every time the site is built, so it cannot quietly go out of date.
Open the complete evidence audit
| Decision | What we bet | Verdict |
|---|---|---|
| Dialogue tutoring as the core modality | A tutor in dialogue beats content alone | Supported — with the 2-sigma correction below |
| Socratic method with ladders + concession (FR-M.3, P1) | Guided questioning, never unguided discovery, answer available on concession | Supported — the guidance caveat lands in our favor |
| Spaced repetition via FSRS | Scheduled retrieval beats restudy | Strongly supported — the two highest-utility techniques known |
| Memory model: strength + read-time fading, half-lives (D-048) | Recall decays exponentially per-item; evidence moves strength | Supported — independently the same model Duolingo derived from data |
| P2: provoke consolidation, never perform it | No auto-summaries; notes/highlights are assessment surfaces | Supported — and more strongly than we knew |
| P3: honesty about mastery | Accurate feedback even when unwelcome | Supported — overconfidence measurably produces underachievement |
| P1: mode choice is the user's | Learner autonomy, never guilt-framed | Supported (self-determination theory) — with a novice caveat we already handle |
| Learner targets Familiar/Working/Deep (D-059) | Ambition is learner-set on one honest scale | Supported (SOLO taxonomy; calibration research) |
| Tutor persona, conversational register (FR-M.16) | A named, polite, conversational tutor that discloses it is software | Supported (personalization + politeness effects) — disclosure is ethics, not efficacy |
| Concept graph + cross-book/mode mastery (D-029, D-059) | Knowledge components independent of where taught | Supported (KLI framework) |
| Book/cards/tutor split; no exercises in the book (P2, §0) | Different learning processes need different instruments | Supported (KLI's three process classes) |
| Verification before drilling (R3) | Never let spaced repetition rehearse an error | Supported by implication — the same memory research that makes FSRS work makes errors durable too |
| Distress exit, never a teaching turn (FR-M.19) | Crisis disclosure exits the frame and signposts humans | Practice standard — crisis-response norms, not RCTs; we say so |
| Gated domains (D-039) | No personal medical/legal/financial/mental-health advice | Practice standard — professional-boundary norms; not an empirical question |
| Retention over engagement metrics (D-046, non-goals) | Measure learning and return, never session count | Supported in principle — engagement ≠ learning is the field's own critique |
| Retrieval by structure, not embeddings (D-007) | Navigable section indexes over similarity search | Engineering judgment — grounding reduces hallucination is supported; the specific mechanism is ours to prove at D.4 |
| Retention-timed reminders (D-065) | Notify from memory state, never streaks | Supported — spacing works only if reviews happen near the optimum; the anti-guilt framing is SDT-consistent |
| Card standard: successive-relearning graduation (D-065) | Durable = recalls across spaced sessions | Supported — Rawson & Dunlosky, incl. in statistics courses specifically |
| Card standard: competitive MC distractors from misconceptions + mandatory feedback (D-065) | Traps become distractors; feedback neutralizes lures | Supported — Little & Bjork on competitive alternatives; Butler/Roediger-line feedback research on lure neutralization |
| Edition targets: readability bands as warning-tier only (D-065) | Formulas inform, humans decide | Supported with the field's own caveat — Flesch-Kincaid is explicitly not a simplification metric; cohesion measures do better; review stays the bar |
| Simulated learners measure the outcome, paired against a read-only control (D-072) | Whether teaching transferred is measurable before real learners exist, but only as a comparison | Engineering judgment, honestly marked — the *comparison* design is sound and mirrors the tutored-vs-read contrast the retention literature uses; what is unproven is whether a role-played learner's gain predicts a human's. Treated as a fast rehearsal of the 0.20 measurement, never as a substitute for it |
| A stopping rule with a "no" branch (D-073) | Tuning must be able to conclude the loop does not teach | Practice standard — pre-registration and stopping rules are standard experimental discipline precisely because a threshold chosen after seeing the data matches the data; not an empirical claim about learning |
| Classroom as a reading surface, not a chat surface (D-086) | A lesson typeset like the book, with the book one tap away, keeps the dialogue grounded and the tutor's turns readable | Design hypothesis, honestly marked — no controlled study compares document-style and bubble-style presentation of tutor dialogue on learning; Mayer's segmenting (learner-paced chunks) is compatible with both, and ICAP is about the learner's activity, not the layout. What is grounded is the coherence with P4 and the reader's layered design (D-056). It ships as a hypothesis with a kill criterion — a 30-second first-user comprehension check and a logged layout toggle — and a change of verdict is a styling change, not an architectural one |
| Capability-typed evidence: recognition, recall, explanation, application and transfer as dimensions of the diary rather than of the score (D-133) | Different task kinds reveal different knowledge, so one number folds them honestly only when the kinds are recorded | Supported — the KLI framework treats these as distinct knowledge components, and transfer routinely dissociates from retention in the literature; our own honesty rule sits on top, since a capability is displayed only once its evidence clears a floor |
| Mastery as a calibrated prediction, scored against later outcomes (D-133) | Strength and confidence together estimate the chance of unassisted success now, and the fold's constants are right exactly insofar as that estimate is calibrated against 30-day results | Supported — calibration against future performance is the standard validity test in knowledge tracing, and predicting observable behaviour dissolves the objection that mastery measures an invisible substance; below bucketable volume the harness must report insufficient data rather than a reliability table |
| Per-learner misconception tracking, from a closed vocabulary (D-133) | What a learner believes incorrectly steers teaching better than scores alone, and refutation works best aimed at the recorded belief | Supported — the refutation-text and misconception-diagnosis literatures already carry D-057's Trap unit and D-065's competitive distractors, and both presuppose exactly this record; what remains ours to prove is how much a closed vocabulary catches |
| Intervention selection by evidence, with the structure disposing (D-133) | Choosing the next teaching move from what the learner has shown beats always descending the same ladder, provided the choice stays verifiable and never overrides what the learner chose | Supported in principle — adaptive-instruction reviews favour evidence-driven variation, and the worked-example and expertise-reversal effects are the canonical cases; the movement criteria are written before the data, and P1 is untouched, because the system varies how it teaches and never what the learner chose |
| Diagnostic placement offered as a suggestion (D-133) | Ten minutes of adaptive prior-knowledge assessment seeds the map and improves the first lesson, while the learner keeps the choice | Supported — prior knowledge is among the strongest predictors of learning, and placement-suggests aligns with self-determination theory the same way D-065's routing rule does; the deterministic walk down the prerequisite graph is engineering judgment, trading a little precision for verifiability and a floor model's economics |