Math Challenge
More

Cognitive Load Theory and Worked Examples in Mathematics

mc-04 · Published: · by Math Challenge Research · 2,435 words · 16 cited sources

Executive summary

256 words

This document was written in English. It is published here in full, unedited.

Verification status

This document carries no [unverified] flag. Every claim in it is tied to a numbered source below.

[unverified] means the claim is stated in the research but was not confirmed against a primary source in the session that produced it. It is published rather than removed, because a research corpus that hides its gaps is not verifiable.

How this research was produced

The 47 documents were produced on 2026-07-31 by independent agents, each instructed not to invent citations and to flag as [unverified] anything it could not confirm against a primary source. The session's web-search quota ran out mid-way, and later agents worked by direct fetch against primary sources. Several sites (ftc.gov, ico.org.uk) block automated fetching, which is why certain legal claims are flagged on purpose.

Findings

1. Theoretical core: working memory, schemas, and three loads

CLT models human cognition per Geary’s split of biologically primary knowledge (evolutionarily prepared, e.g. spoken language) versus biologically secondary knowledge (culturally important but not prepared for, e.g. written arithmetic and algebra) [13]. Because math is biologically secondary it must be explicitly taught, and it is bottlenecked by working memory, which holds only a few novel elements at once and loses them within seconds unless organized into a long-term-memory schema [1][13]. CLT splits load into intrinsic (unavoidable complexity, driven by element interactivity — how many interacting pieces must be held in mind at once), extraneous (added by poor design, no learning value), and germane (effortful processing that builds the schema). Design should minimize extraneous load so spare capacity serves germane processing [1][7][13].

2. The worked-example effect

Sweller and Cooper (1985) taught algebra across five experiments and found that, holding study time constant, worked-example students solved post-test problems roughly twice as fast with about one-fifth the errors of students who solved unaided from the start [1][2]. Unaided problem-solving forces a search process that competes with schema-building; reading an example instead directs full capacity at recognizing solution structure. Sweller calls it “the best known and most widely studied of the cognitive load effects” [1]. Later work (Van Gog, Kester & Paas, 2011) shows example-only and alternating example-problem pairs beating pure problem-solving on other procedural domains (e.g., circuit analysis), so the effect generalizes beyond algebra [1].

3. Expertise reversal and redundancy

The benefit is not permanent. Kalyuga showed across engineering, relay-circuit, and PLC training that worked-examples’ advantage over problem-solving shrinks and reverses as trainees gain experience [3][4]. Mechanism: an efficient schema already exists, so re-showing every step forces processing of unneeded information, wasting capacity that could go to retrieval practice — and denies the practice itself [3][4]. This pushed Sweller to revise his own 1988 claim that problem-solving should generally be minimized: the right amount of unaided solving is a function of current expertise, not curricular stage [4].

4. Fading, completion problems, and self-explanation

Renkl’s fading: start fully worked, then omit one step (a completion problem), then more, until the learner solves unaided. Fading can go forward or backward; backward (omit the last step first) is generally recommended since the final step usually anchors the connection to the goal [5]. Faded examples in geometry produced deeper conceptual understanding than blocks of unfaded examples [5]. Fading alone reliably helps near-transfer but not far-transfer. Atkinson, Renkl and Merrill (2003, J. Educational Psychology, 95(4), 774–783) paired fading with self-explanation prompts — brief questions asking learners to state the principle behind each step, building on Chi’s finding that strong learners spontaneously self-explain and weak learners do not [6][14]. Across two experiments this combination produced medium-to-large gains on both near and far transfer with no added study time — “highly recommendable” per the authors [6]. A broader meta-analysis of self-explanation reports a mean effect size around g = 0.66 across 69 comparisons [14].

5. Split-attention and redundancy in design

Two design-failure effects. Split-attention: understanding requires mentally integrating physically or temporally separated sources (diagram apart from caption, or a step shown before its explanation); integrated, well-designed materials outperform split ones [7]. Redundancy: presenting identical information twice in different channels (e.g., narrating on-screen text word-for-word) adds reconciliation cost with no benefit [7]. For math UI: a solution step and the reasoning behind it should occupy the same visual region at the same time; pick one channel per unit of information.

6. The goal-free effect

A specific-goal problem (“solve for x”) triggers means-ends analysis: holding goal state, current state, their difference, and candidate operators simultaneously — high load, useful for finding an answer but poor for learning method, since attention chases the goal rather than the structure [8]. A goal-free version (“find as many values as you can”) removes the target, so learners work forward opportunistically from what they know, lowering load and directing attention to structure, improving learning even though it doesn’t look like “solving the problem” [8].

7. Critiques, measurement problems, replication debate

CLT’s own leading proponent acknowledges a history of “replication crises and incorporation of other theories,” revised more than once after failed replications — most visibly the 1988 minimize-problem-solving stance, undone by expertise reversal [3][4][9]. The deeper unresolved criticism is measurement: there is no validated, reliable instrument for cognitive load itself; the dominant practice is a single self-report effort-rating item, which cannot be checked for reliability and conflates felt effort with what the manipulation actually did to working memory [9]. Many classic effects are inferred from behavioral outcomes (errors, transfer, time) with load treated as unmeasured — reasonable in aggregate but hard to diagnose case-by-case. Critics also note “cognitive load” sometimes functions as a post-hoc label for any theory-consistent result rather than an independently measured construct — a critique shared with psychology’s broader replication-crisis literature, not unique to CLT [9].

8. Timed practice, speed scoring, and math anxiety

Most relevant to a points-for-speed design. Ashcraft’s program shows math anxiety functions as a concurrent secondary task, occupying working memory (the executive component) that would otherwise serve the arithmetic — producing errors and slowdowns that look like a competence deficit but are a resource deficit induced by anxiety, not by the math [11]. Timed testing reliably reveals anxiety-linked gaps on arithmetic that don’t appear on untimed tests of the same content — the clock, not the math, is what’s being measured under pressure for anxious learners [11][12]. Boaler: timed testing onset is, for a meaningful share of students, the origin point of math anxiety itself, and once anxious, working memory is partly consumed by that anxiety, blocking access to facts they otherwise know — a self-reinforcing loop [12]. The evidence isn’t one-sided: some studies find time-matched conditions can increase accuracy, and one found no significant three-way interaction of memory, anxiety, and timing [11] — real but moderated by anxiety level and by whether “timed” means a hard clock or paced availability.

Synthesis: speed is a legitimate marker of automaticity once a schema is consolidated — fast, effortless retrieval is exactly what germane load “paid for” [1][13]. But speed measured before consolidation doesn’t measure understanding; it measures whatever effortful process the learner is substituting, and a clock on top of that adds exactly the extraneous load CLT says to avoid, at the moment working memory should be protected, not taxed [1][7][8][11].

Design implications for Math Challenge

  1. Show a full worked example before a child’s first attempt at a new pattern (never solved, or last N attempts on that skill failed) rather than “struggle first” — matches the worked-example effect for novices [1][2].
  2. Never pair “new pattern” with a countdown clock. Zero-weight speed scoring on first exposures; introduce timing only once the pattern is solved correctly without a worked example present [11][12].
  3. Fade automatically from a mastery signal, not a fixed item count: recent accuracy, hint usage, unaided-step count select the next rung (full example → one blank → two blanks → full problem), per Renkl’s completion-problem sequence [5][6].
  4. Bias fading backward (omit the last step first), since it anchors the goal connection and should be practiced early [5].
  5. Pair every faded step with a one-tap self-explanation prompt (multiple-choice for younger ages, free text/voice for older) — the one intervention shown to add far-transfer at no time cost [6][14].
  6. Design against split-attention and redundancy in the UI: keep a step and its explanation in the same visual region at once; never narrate and display identical text simultaneously [7].
  7. Use goal-free framings for placement/diagnostic items (“find as many values as you can”) to reveal prior knowledge at lower load, switching to goal-specific scoring once mastery assessment begins [8].
  8. Treat expertise reversal as a hard stop on scaffolding: past a mastery threshold, withhold worked examples/hints by default (available only on request), since redundant scaffolding measurably hurts advanced learners [3][4].
  9. Decouple correctness points from speed points as separately tunable channels; default speed-weight near zero until the fade-ladder reaches “full problem, no hints.” Tell parents/teachers explicitly that raising speed weight increases anxiety-linked variance for below-mastery learners [11][12].
  10. Don’t let a global session/streak timer substitute for per-skill mastery gating — a meta-level clock can reintroduce the same anxiety effect even with untimed problems; make streak timers skippable and never block worked-example access.
  11. Log which fading rung a child needed as a mastery signal for teachers/parents, not just correctness — scaffolding-withdrawal shape is itself diagnostic of schema strength [5][6].
  12. Do not over-instrument “cognitive load” directly (e.g., from response-time variance alone); use behaviorally validated proxies — accuracy trajectory across the fade-ladder, self-explanation quality — since even CLT’s own literature has no validated direct load measure to build on [9].

Open questions for the project owner

  1. Should speed-based points be on by default for any age band given the anxiety evidence, or opt-in per parent/teacher account?
  2. What mastery threshold (e.g., N consecutive correct at “full problem” rung) should gate untimed-to-timed transition per skill?
  3. Should self-explanation prompts be mandatory or skippable, given ages ~4-6 may lack the metacognitive language for them?
  4. Does “find as many values as you can” (goal-free framing) translate cleanly across EN/ES/FR/PT/DE without losing its open-endedness?
  5. Should the fading ladder be visible to the child, or invisible/backend-only?

Sources

  1. [Worked-example effect — Wikipedia](
  2. [Sweller, J., & Cooper, G. A. (1985). The Use of Worked Examples as a Substitute for Problem Solving in Learning Algebra. Cognition and Instruction, 2(1), 59–89 — citation record](
  3. [Expertise reversal effect — Wikipedia](
  4. [The "Expertise Reversal Effect" — Cognitive Load Theory (blog, summarizing Kalyuga et al. 2001, 2003, 2007)](
  5. [Exploring the Use of Faded Worked Examples — ERIC](
  6. [Atkinson, R. K., Renkl, A., & Merrill, M. M. (2003). Transitioning From Studying Examples to Solving Problems: Effects of Self-Explanation Prompts and Fading Worked-Out Steps. Journal of Educational Psychology, 95(4), 774–783 — ERIC record](
  7. [Split attention effect — Wikipedia](
  8. [The Goal-Free Effect (Sweller & Ayres) — Semantic Scholar](
  9. [The Development of Cognitive Load Theory: Replication Crises and Incorporation of Other Theories Can Lead to Theory Expansion — Educational Psychology Review (2023)](
  10. [Cognitive load theory: Research that teachers really need to understand — NSW Department of Education (CESE, 2017)](
  11. [Ashcraft, M., & Krause, J. (2007). Working Memory, Math Performance, and Math Anxiety. Psychonomic Bulletin & Review, 14, 243–248](
  12. [Boaler, J. Speed and Time Pressure Blocks Working Memory (Stanford / youcubed, reprint)](
  13. [Sweller, J. (2019 et al.). Cognitive Architecture and Instructional Design: 20 Years Later. Educational Psychology Review](
  14. [Chi, M. T. H., & Leeuw, N. Eliciting Self-Explanations Improves Understanding — Semantic Scholar](
  15. [Does working memory moderate the effect of fading on math performance? Miller-Cotto et al. British Journal of Educational Psychology (2026)](
  16. [Sweller, J. (1988). Cognitive Load During Problem Solving: Effects on Learning. Cognitive Science, 12(2), 257–285 — Wiley Online Library](

Open questions this document leaves for the owner

These are unanswered on purpose. They are listed, not resolved — turning them into a FAQ would mean inventing answers the document does not contain.

One of 51 research documents, 168,346 words in total, counted at build time from the files themselves. Read this document in the repository