Math Challenge
More

What the research actually says about teaching over the internet

mc-35 · Published: · by Math Challenge Research · 3,247 words · 14 cited sources

Executive summary

The evidence is more modest than the usual "online learning works" pitch suggests. The US Department of Education meta-analysis (Means et al., 2010) found purely online instruction statistically equivalent to face-to-face (+0.05, not significant); the real advantage shows up only in blended (+0.35) and instructor-directed (+0.39) instruction — not in independent, self-directed online learning (+0.05, not significant) [1]. Direct warning for a self-paced app: the online medium alone does not produce the effect; structure and feedback do.

edX video data (Guo et al., 2014) shows median engagement plateaus at 6 minutes at most, and tutorial videos get only 2–3 minutes of real attention regardless of length [2]. The "doer effect" (Koedinger et al.) shows practice carries roughly 6x the learning benefit of reading or watching [6]. Bloom's "2 sigma" claim (1984) comes from a small, never-replicated study; the modern tutoring meta-analysis (Nickow, Oreopoulos & Quan, 2020) finds a pooled effect of just 0.37 SD [3][4]. COVID-era data shows math losses exceeded reading losses (−0.20 to −0.27 SD vs. −0.09 to −0.18 SD), with equity gaps widening another 0.10–0.20 SD [12] — a concrete warning about unsupervised remote instruction in math for children. None of these numbers support aggressive marketing claims; all support a design centered on doing, not watching.

209 words

This document was written in English. It is published here in full, unedited.

Verification status

This document carries no [unverified] flag. Every claim in it is tied to a numbered source below.

[unverified] means the claim is stated in the research but was not confirmed against a primary source in the session that produced it. It is published rather than removed, because a research corpus that hides its gaps is not verifiable.

How this research was produced

The 47 documents were produced on 2026-07-31 by independent agents, each instructed not to invent citations and to flag as [unverified] anything it could not confirm against a primary source. The session's web-search quota ran out mid-way, and later agents worked by direct fetch against primary sources. Several sites (ftc.gov, ico.org.uk) block automated fetching, which is why certain legal claims are flagged on purpose.

Findings

1. Online vs. face-to-face: equivalent, not superior — unless it’s blended and instructor-directed

The US Department of Education’s Evaluation of Evidence-Based Practices in Online Learning (Means, Toyama, Murphy, Bakia & Jones, 2010) meta-analyzed 50 effect sizes from controlled or quasi-experimental studies, 1996-2008 [1]. Headline numbers:

The load-bearing finding for a self-paced product: “independent learning” — the condition closest to an unsupervised app — is the one with a non-significant effect (+0.05). The gains researchers actually found came from instructor direction, collaboration, or added time/materials, none of which a purely self-paced app supplies unless deliberately engineered in.

2. MOOCs: low completion is largely the wrong problem to worry about

MOOC completion rates have been low since the format’s 2012 debut — HarvardX/MITx courses averaged roughly 22% completion in their first year [8]. Stanford research identified four learner archetypes with very different intents: “completers” (5–27%), quick “disengaged” dropouts (6–29%), “auditors” who watch but skip assessments (6–9%), and “samplers” who never intended to finish (39–80%) [8]. Reich & Ruipérez-Valiente’s Science paper “The MOOC Pivot” (2019) argues completion rate is a poor success metric precisely because most enrollees never intended full completion, and that MOOC providers pivoted from a founding mission of open access toward serving already-credentialed professionals paying for credentials — the opposite of the “democratize elite education” pitch the format launched with. One concrete data point: Coursera found learners who paid $30–90 completed at much higher rates than those taking a course free [8] — a commitment-device effect, not a content effect.

The implication: stop measuring success as ”% who complete the whole curriculum,” and instead track engaged mastery among returners plus what predicts return visits.

3. Educational video: short, personal, and beats high production value

Guo, Kim & Rubin’s edX study (2014) analyzed 6.9 million video-watching sessions across four MOOCs [2]:

This lines up with Mayer’s cognitive theory of multimedia learning (Mayer, Multimedia Learning, 2nd ed., 2009): personalization (conversational register beats formal), voice (a natural voice beats a flat one), coherence (cut decoration competing for working memory), redundancy (don’t narrate on-screen text verbatim), segmenting (learner-paced small chunks), signaling (highlight what matters). These independently-derived principles match what Guo et al. found from click-stream data: short, personal, well-segmented beats long, polished, and passive [2][11].

4. Doing beats watching, by a wide margin — the doer effect

Koedinger, Kim, Jia, MacLaren & Bier’s 2015 analysis of an Open Learning Initiative MOOC (“Learning is Not a Spectator Sport: Doing is Better than Watching”) found that the amount of practice a student does predicts subsequent performance far better than the amount of reading or watching they do, holding other factors constant [6]. The widely cited figure from this line of work: doing carries roughly six times the estimated learning benefit of reading/watching per unit of study time. This matches the DOE finding that active/interactive learning beat expository delivery [1], and the edX finding that most viewers never finish a video past a few minutes [2]. Read together: video and reading are, at best, a short on-ramp to a practice problem, not the main event.

5. Bloom’s “2 sigma” — a famous claim that does not survive scrutiny at face value

Bloom’s 1984 paper reported one-to-one tutoring with mastery-learning techniques produced, on average, two standard deviations of gain over conventional classroom instruction — the average tutored student scoring above 98% of controls [4]. This is the number most often invoked to justify “AI tutor” products, and it deserves three caveats. First, it came from two small graduate-student dissertations (Anania; Burke), never independently replicated at that scale [4]. Second, the far larger modern meta-analysis of real-world K-12 tutoring (Nickow, Oreopoulos & Quan, 2020, NBER WP 27476) finds a pooled effect of +0.37 SD — over five times smaller — with similar overall effects for reading and math (reading stronger earlier, math stronger in later grades), and teacher/paraprofessional tutors outperforming volunteers and parents [3]. Third, the tutoring-systems literature (VanLehn, 2011) puts credible expert human one-to-one tutoring closer to 0.7-1.0 SD, not 2.0 — large, but not “2 sigma.” Any claim about an AI tutor should cite the 0.3-0.4 SD range current meta-analytic evidence supports, not Bloom’s figure, until we have our own measured effect.

6. Parasocial, character-led instruction for children: real, but earned over years with a fixed curriculum team

Sesame Street is the best-documented case. The 1994 “Recontact Study” followed preschool viewers into adolescence and found, controlling for confounds, higher English/math/science grades, more pleasure reading, and lower aggression among former viewers — stronger for boys [9]. A 1995 University of Kansas evaluation found disadvantaged children learned as much per viewing hour as advantaged children, though they watched less overall [9]. ETS evaluations (1970-71) found regular 3-year-old viewers outperformed 5-year-old non-viewers on letter recognition [9]. Kearney & Levine (2019, American Economic Review), using the show’s staggered 1969 broadcast-signal rollout as a natural experiment, found exposed children 14% more likely to be at grade-appropriate level in middle/high school, with better adult employment/wage outcomes [9][10]. These results reflect decades of curriculum-research collaboration (the Children’s Television Workshop model), not the character alone — the character is the delivery vehicle for a research-validated curriculum, not a substitute for one.

7. COVID-era remote learning: a cautionary case study, not a validation

Kuhfeld, Soland & Lewis’s analysis of 5.4 million US students in grades 3-8 found fall 2021 math scores 0.20-0.27 SD below same-grade fall-2019 peers, with smaller reading losses (0.09-0.18 SD); poverty-based achievement gaps widened another 0.10-0.20 SD, concentrated in 2020-21 [12]. Separately, research using district-level data (e.g., Goldhaber et al., AERA Open, 2022) found losses were larger where remote/hybrid instructional time was greater relative to in-person time — the amount of remote time tracked the size of the loss. Math suffered most in the least-supervised settings, and disadvantaged students suffered disproportionately. This is a direct caution for an unsupervised, self-paced math app for children: math specifically degrades without structure and supervision when delivered remotely to kids.

Effect size table

Intervention / comparisonEffect sizeSourceApplicability to us
Purely online vs. face-to-face instruction+0.05 (n.s., p=.46)Means et al. 2010, DOE meta-analysis [1]High — closest analogue to unsupervised self-paced app; do not claim superiority from medium alone
Blended vs. face-to-face instruction+0.35 (p<.001)Means et al. 2010 [1]Medium — advantage tied to added time/materials, not to being “online”
Instructor-directed online instruction+0.39 (significant)Means et al. 2010 [1]High — argues for structured guidance, not free exploration
Independent/self-directed online learning+0.05 (not significant)Means et al. 2010 [1]High — this is the condition we are closest to by default
Collaborative online instruction+0.25 (significant)Means et al. 2010 [1]Medium — relevant only if we add social/peer features
Curriculum/pedagogy held identical across conditions+0.13 vs. +0.40 when conditions differMeans et al. 2010 [1]High — most of the “online works” effect is really “more content/time works”
Pooled effect of tutoring programs (all modalities, K-12)+0.37 SDNickow, Oreopoulos & Quan 2020, NBER [3]High — the credible modern baseline for any “AI tutor” claim
Bloom’s one-to-one mastery tutoring vs. classroom (1984, small/unreplicated)~+2.0 SDBloom 1984 [4]Low — historic, contested, do not cite as an expectation
Human expert tutoring vs. no tutoring (contemporary estimate)~+0.7-1.0 SDVanLehn 2011 line of research [5]Medium — ceiling for what live human tutoring plausibly achieves today
Practice (“doing”) vs. reading/watching, per-unit learning benefit~6xKoedinger et al. 2015, “doer effect” [6]Very high — the single most actionable number for session design
MOOC completion rate, elite platforms, 2012~22% averageHarvardX/MITx data [8]Medium — reframes what “success” should be measured as
Math achievement loss, fall 2021 vs. fall 2019 (grades 3-8)−0.20 to −0.27 SDKuhfeld, Soland & Lewis 2022 [12]High — quantifies risk of unsupervised remote math instruction
Reading achievement loss, same period−0.09 to −0.18 SDKuhfeld, Soland & Lewis 2022 [12]Medium — math is the more fragile subject, our subject
Poverty-based achievement gap growth, same period+0.10 to +0.20 SD widerKuhfeld, Soland & Lewis 2022 [12]High — equity risk of self-paced-only delivery
Sesame Street viewers, grade-appropriate attainment in adolescence+14 percentage pointsKearney & Levine 2019 [10]Medium — long-horizon outcome of character-led content, requires years of curriculum rigor to earn

Design implications

  1. Doing must dominate screen time — a reasonable target is 80%+ of session time solving problems, not watching/reading, given the ~6x doer effect [6] and DOE’s active-beats-expository finding [1].
  2. Cap any explanatory video/animation at well under 6 minutes — engagement plateaus at 6 minutes regardless of length, and worked-example content is only really watched for 2-3 minutes [2].
  3. Skip production polish. The costly TV-studio course was less engaging than an informal talking head [2]; spend the budget on instructional design, not a studio.
  4. If we use a tutor character, show a face, not slides-only — 46% vs. 33% post-video engagement for talking-head vs. slides [2], consistent with Mayer’s personalization/voice principles [11].
  5. Apply Mayer’s principles to any tutor dialogue/explanation UI: conversational tone, natural voice, minimal decoration, no verbatim on-screen/narration duplication, learner-paced small chunks, visual highlighting of what matters [11].
  6. Engineer in what the +0.05 “independent learning” number is missing [1]: sequencing that mimics instructor direction (+0.39), and feedback loops — presenting content and letting the student click through will not, by itself, produce a real effect.
  7. Build self-explanation/reflection prompts (e.g., “why was that wrong?”, a confidence check before reveal) — learner reflection/self-monitoring is one of only two practice-level variables the DOE meta-analysis found to matter [1].
  8. Don’t expect generic social/group features to raise learning alone. Group guidance changed how online groups interacted, not how much they learned [1]; any multiplayer feature needs its own measured justification.
  9. Be honest that K-12 evidence is thin. DOE had only 7 K-12 contrasts and warned against generalizing from adult studies [1]; the video and tutoring evidence is likewise mostly adult/mixed-age [2][3]. Decisions for children need our own instrumented data.
  10. Treat math as the more fragile subject in unsupervised settings, not the safer one — COVID math losses (−0.20 to −0.27 SD) outpaced reading (−0.09 to −0.18 SD), worse in more-remote, more-disadvantaged settings [12]. This argues for active structure and monitoring, not “any math practice online is fine.”
  11. Source any outcome marketing claim to the 0.3-0.4 SD tutoring range [3], never to “Bloom’s 2 sigma” [4] — the credible modern ceiling for expert human tutoring is ~0.7-1.0 SD [5], not 2.0, until we have our own measured effect.
  12. Retire ”% who complete the whole curriculum” as the top metric. Most MOOC non-completers never intended to finish [8]; track mastery gain among engaged sessions and return/re-engagement rate instead, watching for early-session drop-off specifically.
  13. Consider a deliberate commitment mechanism (streak, parent-visible report, a diagnostic score to improve) — paid MOOC learners completed at much higher rates than free ones [8], a stakes effect, not a content effect.
  14. A persistent character/mascot is a multi-year curriculum investment, not a one-off asset — Sesame Street’s outcomes came from decades of curriculum research paired with the character [9][10]; a mascot without that underlying rigor won’t reproduce them.
  15. Measure whether the app teaches anything with pre/post + delayed retention, not an immediate post-lesson quiz (which measures working memory, not learning): held-out items, a retention check days-to-weeks later, and a control or staggered rollout where feasible [1].
  16. Track “minutes of actual problem-solving attempted” as the leading internal metric — dosage of doing, not exposure to content, is what the effect-size literature consistently rewards [1][3][6].

Open questions for the project owner

  1. Do we commit engineering effort now to push the product toward the “instructor-directed” evidence tier (+0.39), given that unstructured self-paced use sits in the weakest tier (+0.05, n.s.)?
  2. What commitment/stake mechanism (streak, parent-visible report, diagnostic re-test) are we willing to build, given that stakes strongly predicted MOOC completion?
  3. Do we want a persistent tutor character, given Sesame-Street-level effectiveness required decades of paired curriculum research — is our curriculum team resourced for that, or should we use a plainer UI voice?
  4. What outcome claim, if any, will marketing make, and can we commit to sourcing it to the ~0.3-0.4 SD tutoring range rather than “Bloom’s 2 sigma”?
  5. Will we run our own pre/post + delayed-retention study before making any learning-outcome claim, and on what timeline?
  6. Given math’s larger COVID-era losses vs. reading, should unsupervised use by young children require a parent/guardian setup step or explicit “needs supervision” messaging?
  7. Should any video/animated explanation be capped at a hard duration limit (e.g., 3 minutes) as design policy?

Sources

  1. Means, Toyama, Murphy, Bakia & Jones (2010). Evaluation of Evidence-Based Practices in Online Learning: A Meta-Analysis and Review of Online Learning Studies. U.S. Department of Education
  2. Guo, Kim & Rubin (2014). How Video Production Affects Student Engagement: An Empirical Study of MOOC Videos. L@S 2014
  3. Nickow, Oreopoulos & Quan (2020). The Impressive Effects of Tutoring on PreK-12 Learning: A Systematic Review and Meta-Analysis. NBER Working Paper No. 27476
  4. Bloom, B. S. (1984). The 2 Sigma Problem. Educational Researcher, 13(6). Summary consulted
  5. VanLehn, K. (2011). The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist, 46(4), 197-221
  6. Koedinger, Kim, Jia, MacLaren & Bier (2015). Learning Is Not a Spectator Sport: Doing Is Better Than Watching for Learning from a MOOC. L@S 2015
  7. Reich, J., & Ruipérez-Valiente, J. A. (2019). The MOOC Pivot. Science, 363(6423), 130-131
  8. Wikipedia. Massive open online course — completion-rate data (HarvardX/MITx 2012; Stanford learner-type study; Coursera paid-vs-free completion)
  9. Wikipedia. Sesame Street research — Recontact Study (1994), U. Kansas evaluation (1995), ETS evaluations (1970-71)
  10. Kearney, M. S., & Levine, P. B. (2019). Early Childhood Education by MOOC: Lessons from Sesame Street. American Economic Review, 109(11), 3596-3632
  11. Mayer, R. E. (2009). Multimedia Learning (2nd ed.). Cambridge University Press
  12. Kuhfeld, Soland & Lewis (2022). Test Score Patterns Across Three COVID-19-Impacted School Years. EdWorkingPaper
  13. Goldhaber, Kane, McEachin, Morton, Patterson & Staiger (2022). The Consequences of Remote and Hybrid Instruction During the Pandemic. AERA Open
  14. Fahle, Kane, Reardon & Staiger — Education Recovery Scorecard, Harvard CEPR / Stanford Educational Opportunity Project

Open questions this document leaves for the owner

These are unanswered on purpose. They are listed, not resolved — turning them into a FAQ would mean inventing answers the document does not contain.

One of 51 research documents, 168,346 words in total, counted at build time from the files themselves. Read this document in the repository