How Shello teaches

Learning that lasts is work you do yourself. You have to think the idea through, come back to it as it fades, and know honestly what you can't yet do. Most learning apps are built for the opposite: an answer now, a summary written for you, a score that flatters you. Shello is built around the three harder things. This page says why.

The learning path

Three things happen, in order

  1. The tutor teaches from a real book, and asks before it tells

    So a lesson opens with a question about a chapter you can read yourself, and the evidence for starting there is unusually strong. Researchers compare teaching methods on one scale, called effect size. On it, one-to-one human tutoring scores about d ≈ 0.79 — one of the largest effects known, and smaller than the figure usually quoted. Software that tutors step by step scores about the same as a human tutor.

    In a 2025 randomized trial, students learned more, in less time, from an AI tutor built to enforce good teaching than in an active-learning class. The design did the work. That is Shello's bet too: the teaching lives in how the system is built, never in how clever the model is. Open chat with a model is not what the trial tested.

  2. What you learn is then scheduled against forgetting

    Understanding something once is not the same as still knowing it a month later. Two study techniques stand clear of the rest in the research: testing yourself, and spreading those tests out over time. Shello's cards do both, so each concept comes back as a question around the time you would otherwise lose it. The same research cuts the other way: spaced practice makes whatever you rehearse stick, right or wrong. That is why nothing goes on a card until it has been checked against the chapter it came from.

  3. The one place we won't bend: honesty about what you know

    The mastery map — the record of what you can and can't yet do — has to be honest even when that's unwelcome. Learners who overrate what they know measurably learn and retain less: they ignore feedback and skip the study that would have fixed it. So the map shows what your answers show, and a weak concept looks weak. Everything else — pace, mode, how far you want to go — is yours to set, and we never use guilt to steer you.

See it in practice

A lesson begins before the answer

Tutor waits for your reasoning, then uses the chapter to test and refine it. The question is not a delay before teaching; it is where the teaching begins.

A value far above the rest is added to a small data set. Before calculating, what do you expect it to do to the mean — and why?

Two myths from our own field that we won't repeat

  1. The first is the figure usually quoted for tutoring. Bloom's famous "2-sigma" claim — that one-to-one tutoring lifts students by 2 standard deviations — is folklore in its strong form. The reviewed figure is d ≈ 0.79; the myth inflates it about 2.5 times. We build on the honest number.

  2. The second is about how people study. Summarizing, highlighting, and rereading rate among the least useful techniques in the research, and they are the ones most learning apps automate. That is why Shello has no auto-summary button. The book gets you to pull the ideas together yourself; it never does it for you.

The evidence, decision by decision

Everything above is the short version. Below is the whole audit: each decision, what we bet on it, and the verdict. "Supported" means published studies back the choice; "practice standard" means professional norms do, where no controlled study exists or applies. The audit is a document kept with our code. The table is rebuilt from it every time the site is built, so it cannot quietly go out of date.

Open the complete evidence audit
DecisionWhat we betVerdict
Dialogue tutoring as the core modalityA tutor in dialogue beats content aloneSupported — with the 2-sigma correction below
Socratic method with ladders + concession (FR-M.3, P1)Guided questioning, never unguided discovery, answer available on concessionSupported — the guidance caveat lands in our favor
Spaced repetition via FSRSScheduled retrieval beats restudyStrongly supported — the two highest-utility techniques known
Memory model: strength + read-time fading, half-lives (D-048)Recall decays exponentially per-item; evidence moves strengthSupported — independently the same model Duolingo derived from data
P2: provoke consolidation, never perform itNo auto-summaries; notes/highlights are assessment surfacesSupported — and more strongly than we knew
P3: honesty about masteryAccurate feedback even when unwelcomeSupported — overconfidence measurably produces underachievement
P1: mode choice is the user'sLearner autonomy, never guilt-framedSupported (self-determination theory) — with a novice caveat we already handle
Learner targets Familiar/Working/Deep (D-059)Ambition is learner-set on one honest scaleSupported (SOLO taxonomy; calibration research)
Tutor persona, conversational register (FR-M.16)A named, polite, conversational tutor that discloses it is softwareSupported (personalization + politeness effects) — disclosure is ethics, not efficacy
Concept graph + cross-book/mode mastery (D-029, D-059)Knowledge components independent of where taughtSupported (KLI framework)
Book/cards/tutor split; no exercises in the book (P2, §0)Different learning processes need different instrumentsSupported (KLI's three process classes)
Verification before drilling (R3)Never let spaced repetition rehearse an errorSupported by implication — the same memory research that makes FSRS work makes errors durable too
Distress exit, never a teaching turn (FR-M.19)Crisis disclosure exits the frame and signposts humansPractice standard — crisis-response norms, not RCTs; we say so
Gated domains (D-039)No personal medical/legal/financial/mental-health advicePractice standard — professional-boundary norms; not an empirical question
Retention over engagement metrics (D-046, non-goals)Measure learning and return, never session countSupported in principle — engagement ≠ learning is the field's own critique
Retrieval by structure, not embeddings (D-007)Navigable section indexes over similarity searchEngineering judgment — grounding reduces hallucination is supported; the specific mechanism is ours to prove at D.4
Retention-timed reminders (D-065)Notify from memory state, never streaksSupported — spacing works only if reviews happen near the optimum; the anti-guilt framing is SDT-consistent
Card standard: successive-relearning graduation (D-065)Durable = recalls across spaced sessionsSupported — Rawson & Dunlosky, incl. in statistics courses specifically
Card standard: competitive MC distractors from misconceptions + mandatory feedback (D-065)Traps become distractors; feedback neutralizes luresSupported — Little & Bjork on competitive alternatives; Butler/Roediger-line feedback research on lure neutralization
Edition targets: readability bands as warning-tier only (D-065)Formulas inform, humans decideSupported with the field's own caveat — Flesch-Kincaid is explicitly not a simplification metric; cohesion measures do better; review stays the bar
Simulated learners measure the outcome, paired against a read-only control (D-072)Whether teaching transferred is measurable before real learners exist, but only as a comparisonEngineering judgment, honestly marked — the *comparison* design is sound and mirrors the tutored-vs-read contrast the retention literature uses; what is unproven is whether a role-played learner's gain predicts a human's. Treated as a fast rehearsal of the 0.20 measurement, never as a substitute for it
A stopping rule with a "no" branch (D-073)Tuning must be able to conclude the loop does not teachPractice standard — pre-registration and stopping rules are standard experimental discipline precisely because a threshold chosen after seeing the data matches the data; not an empirical claim about learning
Classroom as a reading surface, not a chat surface (D-086)A lesson typeset like the book, with the book one tap away, keeps the dialogue grounded and the tutor's turns readableDesign hypothesis, honestly marked — no controlled study compares document-style and bubble-style presentation of tutor dialogue on learning; Mayer's segmenting (learner-paced chunks) is compatible with both, and ICAP is about the learner's activity, not the layout. What is grounded is the coherence with P4 and the reader's layered design (D-056). It ships as a hypothesis with a kill criterion — a 30-second first-user comprehension check and a logged layout toggle — and a change of verdict is a styling change, not an architectural one
Capability-typed evidence: recognition, recall, explanation, application and transfer as dimensions of the diary rather than of the score (D-133)Different task kinds reveal different knowledge, so one number folds them honestly only when the kinds are recordedSupported — the KLI framework treats these as distinct knowledge components, and transfer routinely dissociates from retention in the literature; our own honesty rule sits on top, since a capability is displayed only once its evidence clears a floor
Mastery as a calibrated prediction, scored against later outcomes (D-133)Strength and confidence together estimate the chance of unassisted success now, and the fold's constants are right exactly insofar as that estimate is calibrated against 30-day resultsSupported — calibration against future performance is the standard validity test in knowledge tracing, and predicting observable behaviour dissolves the objection that mastery measures an invisible substance; below bucketable volume the harness must report insufficient data rather than a reliability table
Per-learner misconception tracking, from a closed vocabulary (D-133)What a learner believes incorrectly steers teaching better than scores alone, and refutation works best aimed at the recorded beliefSupported — the refutation-text and misconception-diagnosis literatures already carry D-057's Trap unit and D-065's competitive distractors, and both presuppose exactly this record; what remains ours to prove is how much a closed vocabulary catches
Intervention selection by evidence, with the structure disposing (D-133)Choosing the next teaching move from what the learner has shown beats always descending the same ladder, provided the choice stays verifiable and never overrides what the learner choseSupported in principle — adaptive-instruction reviews favour evidence-driven variation, and the worked-example and expertise-reversal effects are the canonical cases; the movement criteria are written before the data, and P1 is untouched, because the system varies how it teaches and never what the learner chose
Diagnostic placement offered as a suggestion (D-133)Ten minutes of adaptive prior-knowledge assessment seeds the map and improves the first lesson, while the learner keeps the choiceSupported — prior knowledge is among the strongest predictors of learning, and placement-suggests aligns with self-determination theory the same way D-065's routing rule does; the deterministic walk down the prerequisite graph is engineering judgment, trading a little precision for verifiability and a floor model's economics

Choose something worth learning