Math Challenge
More

Designing Engaging, Pedagogically Sound Math Challenges: Rich Task Traditions, Item Formats, and Standards

mc-36 · Published: · by Math Challenge Research · 5,713 words · 20 cited sources

Executive summary

439 words

This document was written in English. It is published here in full, unedited.

Verification status

This document carries no [unverified] flag. Every claim in it is tied to a numbered source below.

[unverified] means the claim is stated in the research but was not confirmed against a primary source in the session that produced it. It is published rather than removed, because a research corpus that hides its gaps is not verifiable.

How this research was produced

The 47 documents were produced on 2026-07-31 by independent agents, each instructed not to invent citations and to flag as [unverified] anything it could not confirm against a primary source. The session's web-search quota ran out mid-way, and later agents worked by direct fetch against primary sources. Several sites (ftc.gov, ico.org.uk) block automated fetching, which is why certain legal claims are flagged on purpose.

Findings

1. Rich task design: low floor, high ceiling, wide walls

The phrase “low floor, high ceiling” originates with Seymour Papert’s constructionist work and was extended with “wide walls” by Mitchel Resnick to describe creative tools (like Scratch) that let anyone begin easily, allow experts to go far, and support many different paths and styles of work rather than one correct route. NRICH (University of Cambridge) adopted this triad as its working definition of a “rich task”: low floor so every student has an entry point, high ceiling so advanced students are still stretched, and wide walls so the task supports multiple representations and solution strategies, not just multiple final answers [20]. This is distinct from, and stronger than, “differentiated difficulty” — wide walls specifically means the mathematics itself branches (e.g., a task solvable by drawing, by a table, by algebra, or by a recursive argument), not just that easier and harder versions of the same procedure exist.

A complementary, more mechanically checkable framework layers on top: Stein and Smith’s cognitive-demand levels distinguish “procedures with connections” (a procedure used to build meaning, connected to underlying concepts, represented multiple ways, requiring some effort) from “doing mathematics” (non-algorithmic, exploratory, requiring self-monitoring, with real ambiguity in the solution path) — and Cohen and colleagues’ “group-worthy task” criteria add: centers on an important mathematical idea, has multiple entry points and strategies, requires explanation/justification, and minimizes non-mathematical noise that would let a student succeed without engaging the math [20]. Together these give Math Challenge a rubric that can be applied to any drafted item, independent of topic: does it have a genuine entry point for a struggling student, genuine room to extend for a strong one, and more than one legitimate way to attack it.

2. Three-act math and the pseudocontext problem

Dan Meyer’s “three-act math” structure responds to a specific failure mode of traditional word problems: math bolted onto a scenario the student never actually needed math to understand (“pseudocontext”). Act One presents an intriguing image or short video with no numbers and no explicit question — the goal is that the student’s own curiosity generates the question. Act Two supplies only the data or tools the students themselves identify as necessary, so the mathematical work is purposeful rather than imposed. Act Three reveals the real outcome, letting students check their reasoning against reality rather than an answer key [1]. Meyer frames the payoff explicitly: after the math, the student should feel more powerful — able to predict or explain something they could not before — rather than having completed an assigned exercise. For Math Challenge, this suggests a small, high-production-cost tier of items (video- or image-led, one Act One per unit) sitting above a much larger tier of cheaper drilled items, rather than trying to three-act every item in a 2,500-item bank.

3. Shell Centre / Malcolm Swan task genres

The Mathematics Assessment Project (Shell Centre, University of Nottingham, directed in its later phase by Malcolm Swan) produced roughly 100 “Classroom Challenge” formative-assessment lessons spanning grades 6 through college-readiness, split broadly into two categories: Concept Development lessons, which reveal and refine understanding of a mathematical idea, and Problem Solving lessons, which target non-routine problems and involve cycles of students producing a solution, receiving structured feedback (often via sample student work), and revising [7][8]. Swan’s own theoretical framing goes further, identifying five distinct purposes mathematics instruction can serve — fluency with facts/skills, interpreting concepts and representations, developing strategies for investigation and problem solving, awareness of the educational system’s values, and appreciating mathematics’ power in society — and argues that task design should follow deliberately from which purpose is intended, because different purposes imply different learning theories and therefore different task shapes [9]. He further describes five task types that specifically support concept development: classifying objects, interpreting multiple representations, evaluating mathematical statements, creating problems, and analyzing reasoning or solutions [9]. This is directly reusable as an item-tagging taxonomy: every Math Challenge item can be labeled not only by topic and difficulty but by which of these five purposes it serves, which in turn should influence its format (fluency items suit quick MCQ/numeric entry; “evaluate a statement” suits true/false-with-justification or spot-the-error; “create a problem” suits an open-response or problem-posing format).

4. Open Middle problems

Robert Kaplinsky’s Open Middle format is defined by three structural properties: a closed beginning (every student starts from the same given information), a closed end (there is one correct final answer, or one optimal answer among several valid ones), and an open middle (there are multiple legitimate paths to get there) [2]. This differs from a fully open-ended task in that it keeps an unambiguous, machine-checkable target answer — which matters enormously for an auto-graded product — while still requiring genuine mathematical reasoning rather than single-step recall, because the “middle” (the search for digit placements, factor choices, or optimal combinations) is where the cognitive work lives. A canonical example: “Using the digits 1–9 at most one time each, fill in the boxes to create the smallest possible product” — trivially auto-gradable (compare submitted digits’ product to the known minimum), but resistant to guessing because there is no shortcut to the answer other than reasoning about place value and factor size. Open Middle problems are optimization-flavored (“find the smallest/largest/all possible”) rather than single-step, and Kaplinsky notes they occupy a middle ground of design effort: more elaborate than a plain fluency drill, but far less elaborate than a full multi-day performance task [2].

5. Short, reusable classroom routines: WODB, Visual Patterns, Number Talks

Three widely-adopted “routine” formats are notable because they are extremely cheap to author in bulk and translate almost directly into digital item templates. Which One Doesn’t Belong (WODB), created by Mary Bourassa in 2014 after adapting an early-childhood classification activity for a calculus class, presents four items (numbers, shapes, graphs, equations) in a grid and asks which one doesn’t belong — deliberately with no single correct answer, since the goal is that each of the four items can be justified as the “odd one out” for a different, valid mathematical reason [3]. Bourassa’s own design heuristic: a strong WODB has three items sharing a characteristic the fourth lacks, for at least one reasonable partition of the four — meaning multiple different partitions should all be defensible [4]. Because WODB has no single answer, it is not naturally auto-gradable in a points sense, but works well as a “select your item and justify” format scored on the quality/validity of the given reason rather than which item was picked, or used as an ungraded warm-up/discussion item type.

Visual Patterns (Fawn Nguyen) presents a sequence of figures (step 1, 2, 3…) that grow by a visible rule, and asks students to predict a distant step (e.g., step 43) or produce a general equation — training exactly the move from concrete pattern-noticing to algebraic generalization [10]. This is highly parameterizable: a single “growth-pattern generator” (linear, quadratic, or geometric growth rules applied to dot/tile arrangements) can mechanically produce large numbers of isomorphic items differing only in the specific rule and requested step number.

Number Talks (widely associated with Ruth Parker, Kathy Richardson, and popularized by Sherry Parrish and Jo Boaler) is a short daily routine centered on solving a single computation mentally, then having students share and compare different valid mental strategies for the same problem — the point is not the answer but the visible variety of correct reasoning paths. This is naturally suited to a “show your strategy” or “match your method to a name” digital format rather than a bare numeric-entry item, since the pedagogical payload is method comparison, echoing Swan’s “evaluate/compare representations” task type [9].

6. Fermi estimation problems

Fermi problems — named for Enrico Fermi’s characteristic ability to produce accurate order-of-magnitude estimates from minimal data, illustrated by his famous on-the-spot estimate of the Trinity test’s explosive yield from how far paper scraps were blown — are estimation tasks with no readily available exact data, solved by decomposing the quantity into a chain of roughly-estimated factors [13]. Because errors in the individual factor estimates tend to partly cancel when multiplied together, the product of several rough guesses is often surprisingly close to the true order of magnitude even when no single guess was very accurate [13]. Pedagogically this trains comfort with uncertainty, dimensional reasoning, and sanity-checking a large calculated result against intuition — and it maps to a specific item format: numeric entry scored correct within an order-of-magnitude (or percentage) tolerance band, not an exact match, which is a fundamentally different grading contract than most school arithmetic.

7. Problem-posing and the “suspension of sense-making” in word problems

A separate strand of research examines the gap between executing a procedure correctly and connecting that procedure to a realistic situation’s meaning. General findings in this area (surveyed in overviews of word-problem pedagogy) note that the comprehension difficulty in word problems is often not that students cannot execute the arithmetic, but that they lack a firm connection between the mathematical operations and the semantics of the realistic scenario the problem describes [14] — students can compute correctly while effectively “switching off” whether the resulting number makes sense in the world the problem describes (the widely cited genre of test items where students divide leftover people into “buses needed” or similar without checking whether the numeric answer is even physically sensible illustrates the same phenomenon). This is the same failure mode three-act math and Open Middle both push against from different angles: three-act math prevents it by making the real-world stakes concrete and student-generated before any number appears [1]; Open Middle prevents it by making the “middle” itself the reasoning task, so there is no procedure to run blindly [2]. For Math Challenge, the actionable version of this research is a design rule: any word problem item should be checkable for whether a “plausible but senseless” wrong answer (negative time, fractional people, an area bigger than the containing shape) is a common enough error to warrant a dedicated “does this answer make sense?” distractor or a post-answer reflection prompt.

8. Anatomy of a good multiple-choice item and misconception-based distractors

Standard item-writing guidance treats distractor construction as deliberate engineering: distractors are written to be “plausible yet clearly incorrect,” and a student’s initial pull toward a wrong option should come from surface plausibility the item writer built in on purpose, not from an implausible option nobody would pick (e.g., “Detroit” as a distractor in a question about Indian cities is a textbook example of a bad, non-plausible distractor). A known weakness of the format is “test-wiseness” — clues in option construction (length, grammatical fit, absolute qualifiers) that let a test-savvy student score above their actual knowledge, which is a reason to keep distractor length/style/grammar uniform across options. The strongest, most defensible distractors are not randomly wrong numbers but the actual numeric or symbolic output of a documented misconception — e.g., for “-2 - (-5)”, offering “-7” (sign-of-subtraction error) and “-3” (drops the negation) as distractors, rather than arbitrary wrong numbers — because a distractor built from a real bug is diagnostic (which wrong answer was picked tells you which misconception is present), whereas an arbitrary wrong number only tells you the student erred, not how.

9. Item formats beyond multiple choice

A useful working catalogue for Math Challenge (developed in detail in the Item Format Catalogue table below) includes: numeric entry with a tolerance band (needed for Fermi-style estimation and any answer with legitimate rounding); drag-to-order (sequencing steps of a proof, or ordering numbers/fractions on a line); matching (pairing an expression with its graph, or a term with its definition); hotspot/select-point or click-the-figure (identify a region, angle, or point on a diagram); construct-a-graph (place points, draw a line of given slope, plot a function) — supported natively as graphicGapMatchInteraction, selectPointInteraction, positionObjectInteraction and drawingInteraction in the QTI 3.0 specification [6]; fill-the-blank equation (parametrized gap-fill within a symbolic expression, gapMatchInteraction/textEntryInteraction in QTI terms) [6]; sort-into-categories (classify a list of numbers/shapes into buckets — a natural digitization of Swan’s “classifying objects” task type) [9]; spot-the-error (present a worked solution with an injected bug and ask the student to locate/name it — directly implements IES-style “use solved, including incorrect, problems”); estimation-with-a-range (numeric entry scored by containment in an interval, the natural Fermi-problem format) [13]; multi-select (QTI’s choiceInteraction with maxChoices > 1) [6]; and interactive manipulatives (drag tiles/counters, virtual tangram pieces, virtual base-ten blocks) [16].

10. Puzzle traditions: Japanese, Chinese, Russian, Vedic

Japanese puzzle traditions offer several formats already proven at scale with millions of solvers. Kakuro (“Cross Sums”), despite its Japanese name, was actually invented in 1966 by Canadian Jacob E. Funk at Dell Magazines and later gained its lasting popularity in Japan; it requires filling a grid so that each horizontal/vertical run of white cells sums to a given clue using no repeated digit, exercising combinatorial reasoning about which digit sets can produce a target sum [11]. KenKen, invented in 2004 by Japanese teacher Tetsuya Miyamoto explicitly as an “instruction-free” brain-training tool, combines a Latin-square constraint (each digit once per row/column) with “cage” regions that must produce a target value under a stated arithmetic operation, and is reported in use by over 30,000 U.S. teachers for arithmetic fluency and logical deduction practice [12]. Naoki Inaba’s “area maze” (menseki meiro) puzzles — grids of nested rectangles with some side lengths and areas given, others to be found using only arithmetic and geometric insight, no algebra or calculator — are a further well-known Japanese genre in this same family, though this report was not able to verify a citable primary source for that specific format during this research pass and it is included here as background context rather than a cited finding.

Chinese traditions contribute the tangram (七巧板, “seven boards of skill”), a dissection puzzle of seven flat pieces (five triangles, one square, one parallelogram) that recombine into a square or countless other silhouette figures, used pedagogically to build spatial visualization, congruence, and area-conservation intuition (since all tangram figures share the same total area) [16]. The Nine Chapters on the Mathematical Art, the foundational classical Chinese mathematics text, is organized as nine topic chapters (field areas and fractions, proportional exchange, proportional distribution, root extraction from area/volume, construction volumes, taxation-style proportion problems, systems of two linear equations via “excess and deficit,” general linear systems via Gaussian-elimination-like methods, and right-triangle/Pythagorean problems), each following a strict “problem, then solution, then method” format — a structure directly reusable as an item-authoring template (worked example → generalizable method) [15].

Russian math circles (kruzhok), tracing to Bulgaria before 1907 and formalized in the USSR in the 1930s, are voluntary, non-curricular problem-solving groups built around genuinely wanting to be there and a social context for enjoying mathematics, typically Socratic and exploration-driven rather than lecture-driven, sometimes olympiad-oriented and sometimes deliberately not; the tradition reached the US in 1994 via Robert and Ellen Kaplan at Harvard, brought by émigrés who had themselves attended as teenagers [18]. The math-circle problem genre (self-contained, elegant, requiring insight rather than a memorized procedure) is a strong source for high-ceiling “wide wall” items distinct from curriculum-aligned drill.

Vedic mathematics requires a caveat: the 1965 book by Bharati Krishna Tirtha presents 16 “sutras” (e.g., shortcuts for squaring numbers ending in 5, or specific multiplication patterns) claimed to be drawn from the ancient Atharvaveda, but this attribution is rejected by mathematics historians (S. G. Dani, Kim Plofker, K. S. Shukla), who find the linguistic style is modern Sanskrit, the techniques rely on decimal notation unknown in the Vedic period, and the content has essentially nothing in common with actual historical Vedic-era mathematics [17]. The individual calculation shortcuts themselves can still be legitimate, useful arithmetic tricks worth including as a “speed technique” item category, but any user-facing framing should avoid presenting them as authentic ancient scripture, both because it is factually disputed and because (per the source found) the framing has become politically contested in India’s curriculum debates [17].

11. Parameterized item templates vs. handwritten items

Automatic Item Generation (AIG) is the formal discipline behind turning one hand-authored “item model” into many delivered items. The standard vocabulary: radicals are the structural features of an item model that determine its difficulty (e.g., number of steps, size of numbers, presence of a distractor-triggering feature); incidentals are surface features that vary without changing difficulty (which specific numbers, names, or cover story is used); and isomorphs are the resulting items that share identical radicals but differ only in incidentals [19]. AIG’s practical value is squarely aligned with a 2,500-item MVP target: a test specialist designs one carefully vetted item model with variable slots, and an algorithm fills the slots, producing many parallel items far faster than one-at-a-time hand authoring, while — if the model is sound — keeping difficulty and quality comparable across the generated family [19]. The explicit trade-off is that quality now depends entirely on whether the underlying model correctly separates radicals from incidentals; a poorly designed model can silently generate items that are much harder or easier than intended, or that share an exploitable surface pattern a repeat solver could learn. This argues for building a smaller number (dozens, not thousands) of high-quality item models — one per skill/misconception pairing identified in topic-specific research already gathered for this project — each parameterized to emit many isomorphic instances, rather than treating “2,500 items” as 2,500 independent authoring tasks.

12. The QTI 3.0 standard

QTI (Question and Test Interoperability), maintained by 1EdTech, is the dominant open standard for packaging assessment items, tests, and result data so they move between authoring tools, item banks, and delivery platforms without vendor lock-in [5]. QTI 3.0 consolidates the earlier QTI 2.x line with the APIP accessibility specification and moves to web-friendly markup (custom elements usable as standard web components), while retaining native support for computer-adaptive testing and portable custom interactions [5]. Concretely, QTI 3.0 defines 21 standardized interaction types, directly covering nearly every format in the catalogue above: choiceInteraction (single/multi-select MCQ), textEntryInteraction/inlineChoiceInteraction (short numeric/text or embedded-choice fill-in), extendedTextInteraction (free response), gapMatchInteraction/graphicGapMatchInteraction (fill-the-blank, including onto an image), matchInteraction/associateInteraction (pairing two sets), orderInteraction/graphicOrderInteraction (drag-to-sequence, including regions of an image), hotspotInteraction/selectPointInteraction (click-the-region/point), positionObjectInteraction (drag an object onto a target image), sliderInteraction (bounded numeric selection by dragging), drawingInteraction (freehand mark-up of a canvas), and uploadInteraction (file submission) [6]. Adopting QTI 3.0 item-model naming even informally (regardless of whether Math Challenge implements full QTI XML export in the MVP) gives a vetted, accessibility-aware vocabulary for the item-authoring schema instead of inventing format names from scratch, and keeps a real export path open later without a rewrite.

Item format catalogue

FormatAges it suitsWhat it measuresAuto-gradable?Solver-resistant?Build effort
Multiple choice (single-select), misconception-based distractorsAll agesRecognition + which specific error a wrong pick revealsYes, triviallyLow — guessable, and vulnerable to memorized answer keys unless items are varied/parameterizedLow per item, higher to build a real distractor library from documented errors
Multi-select8+Ability to identify all valid cases, not just oneYesMedium — harder to guess than single-selectLow
Numeric entry with tolerance8+Computation, rounding judgmentYes (range check)MediumLow
Numeric entry, exactAll agesPrecise computationYesLow–mediumLow
Estimation with a range (Fermi-style)10+Order-of-magnitude reasoning, comfort with uncertainty [13]Yes (interval containment)High — no procedure to memorize, only a reasoning chainMedium (needs a believable real-world quantity)
Fill-the-blank equation / expression8+Structural/symbolic understanding, one missing pieceYesMediumLow–medium
Drag-to-order6+Sequencing (proof steps, magnitude ordering, process order)YesMediumMedium (needs a drag UI)
Matching / pairing8+Linking representations (graph↔equation, term↔definition)YesMediumMedium
Hotspot / select-point / click-the-figure6+Geometric identification, reading a diagramYesMediumMedium–high (needs image regions)
Drag-onto-image (position object)8+Spatial placement (plot a point, place a piece)YesMediumMedium–high
Construct-a-graph (plot points/line/function)11+Graphing fluency, function behaviorPartially — needs tolerance logicHighHigh
Sort into categories6+Classification (Swan’s “classifying objects” task type) [9]YesMediumLow–medium
Spot-the-error (worked solution with injected bug)10+Diagnosing a specific misconception, not just producing a correct answerYes if bug location is discrete/taggedHigh — requires understanding the method, not pattern-matching an answerMedium (needs a bug library)
Which One Doesn’t Belong (justify your pick)All agesDiscourse, noticing attributes, multiple valid reasoningNo (graded on justification quality, or ungraded)High (no single “right” pick to look up)Low
Visual pattern → predict step N / write rule8+Generalization, early algebraic reasoning [10]YesMedium–highLow (highly parameterizable)
Number talk (choose/describe your mental strategy)6+Metacognition, strategy comparison, number sensePartially (multiple valid strategies)MediumLow
Open Middle (find the optimum)8+Multi-path reasoning toward one closed answer [2]Yes (compare to known optimum)High — no shortcut around the searchMedium
Three-act math (image/video-led, no numbers first)8+Motivation, question-posing, applying math to a real prediction [1]Partially (final numeric answer, yes; the framing, no)High (context changes every time)High (needs media)
Puzzle formats (KenKen/Kakuro/tangram-style)8+Combinatorial reasoning, constraint satisfaction, spatial reasoning [11][12][16]YesHigh — resists rote memorization by constructionMedium–high (needs a generator/solver)
Interactive manipulatives (virtual base-ten blocks, fraction bars, tangram)3–11Concrete-to-abstract bridgingPartially (depends on task wrapped around it)MediumHigh

Design implications

  1. Adopt “low floor, high ceiling, wide walls” plus the cognitive-demand/group-worthy-task rubric as a mandatory pre-publication checklist for every item template, not just the hand-authored showcase items — even a parameterized MCQ can be scored against “does this have more than one legitimate path” [20].
  2. Tag every item (or item model) with which of Swan’s five instructional purposes it serves (fluency, concept, non-routine problem-solving, mathematical language, application), since that should drive format choice, not just topic [9].
  3. Build the MVP’s 8–10 supported item formats in roughly this order of build-vs-payoff: (1) multiple choice with misconception distractors, (2) numeric entry with tolerance, (3) multi-select, (4) fill-the-blank equation, (5) sort into categories, (6) matching/pairing, (7) drag-to-order, (8) Open Middle-style optimum search, (9) spot-the-error, (10) hotspot/click-the-figure — this front-loads the formats with the lowest build cost and highest auto-gradability, and defers the two most production-heavy formats (construct-a-graph, three-act video-led items) to a later release.
  4. Treat “three-act” style items and true rich/NRICH-style open tasks as a small, deliberately non-scalable premium tier (tens of items, not thousands) rather than trying to force the format across all 2,500 items — the production cost (media, open-ended scoring) is fundamentally incompatible with bulk generation [1][20].
  5. Invest the majority of authoring effort in item models (per Automatic Item Generation’s radicals/incidentals/isomorphs framework), not individual items: pick one misconception per model (sourced from this project’s topic-specific misconception research), define which numeric/contextual features are “radicals” that must be held difficulty-constant, and let an algorithm vary the “incidentals” to reach volume [19].
  6. Build a real, citable distractor library per skill, keyed to documented misconceptions (not randomly generated wrong numbers), and make the MVP’s diagnostic signal be which distractor was picked, not just right/wrong — this is the single highest-leverage design choice for turning a drill bank into an adaptive tutor.
  7. Use Open Middle’s “closed beginning/closed end/open-middle” structure as the default template for any “find the best/smallest/largest” item, since it is simultaneously auto-gradable and resistant to answer lookup — a rare combination [2].
  8. Give Fermi/estimation-with-range its own first-class item type (numeric entry, tolerance-band scoring) rather than shoehorning it into exact-numeric-entry — the grading contract is fundamentally different (interval containment, not equality) [13].
  9. Reserve “spot-the-error” for the specific misconceptions already catalogued in this project’s per-topic research (e.g., the algebra misconception taxonomy), so the injected bug in the worked solution is always a real, documented error, and tag which named bug each spot-the-error item tests.
  10. Adopt QTI 3.0’s interaction-type names (choiceInteraction, textEntryInteraction, gapMatchInteraction, matchInteraction, orderInteraction, hotspotInteraction, positionObjectInteraction, sliderInteraction, drawingInteraction, etc.) as the internal schema vocabulary for the item-authoring data model now, even without implementing QTI XML import/export in the MVP — this keeps a real interoperability/export path open later without a schema rewrite [5][6].
  11. Include a “no single right answer, graded on justification” item class (modeled on Which One Doesn’t Belong) explicitly outside the points/mastery scoring pipeline, since forcing a discourse-based routine into a right/wrong score defeats its purpose [3][4].
  12. Mine the Nine Chapters’ “problem → solution → method” structure and Russian math-circle problems as a source of hand-craftable high-ceiling items distinct from curriculum drill, explicitly labeled as an enrichment/challenge track rather than mixed anonymously into core progression [15][18].
  13. If Vedic-style calculation shortcuts are included as a “speed technique” content category, present them as named arithmetic tricks (e.g., “the base-10 squaring trick”) without a historical-authenticity claim, given the disputed/politicized provenance of the “Vedic mathematics” framing [17].
  14. For every word-problem-style item (any format, not just MCQ), add either a “does this answer make sense” distractor (a numerically-derivable but physically absurd option) or a lightweight post-answer reflection prompt, since mechanical procedure-execution without real-world sense-checking is a documented, recurring failure mode independent of the specific math skill being tested [1][14].

Open questions for the project owner

  1. Should the MVP commit to a QTI 3.0-shaped internal schema from day one (even without XML export), or is that premature standardization the team would rather defer until there’s an actual need to interoperate with an external item bank or LMS?
  2. For the “no single right answer” item class (Which One Doesn’t Belong-style), should Math Challenge score participation/engagement only, use a lightweight AI-graded justification score, or exclude this class from any scoring/streak mechanic entirely?
  3. Given that Open Middle and Fermi-style items are the two formats identified as most resistant to answer-lookup/solver cheating, should the MVP’s early XP/points economy weight those formats more heavily than plain MCQ to discourage students from gravitating toward the “cheapest” item type?
  4. How much of the 2,500-item target should be reserved for the small, high-production “premium” tier (three-act, true NRICH-style rich tasks, curated puzzle-tradition items) versus the bulk parameterized-template tier — is there a target ratio (e.g., 95/5) the owner wants set explicitly now, or should it be discovered empirically after the first content sprint?
  5. Should the misconception-based distractor library be built topic-by-topic as each topic’s research lands (piggybacking on the per-topic misconception research already produced for this project), or centralized as a separate cross-cutting authoring pass after all topic research is in?

Sources

  1. Dan Meyer, "Three-Act Math" (blog category)
  2. Robert Kaplinsky / Open Middle, "What's Open Middle?"
  3. Mary Bourassa, "Announcing the Which One Doesn't Belong? Website!"
  4. Amplify Polypad, "Which One Doesn't Belong" lesson guide
  5. 1EdTech, "QTI (Question & Test Interoperability)" standard overview
  6. 1EdTech / IMS Global, QTI 3.0 Implementation Guide (interaction types)
  7. Mathematics Assessment Project (Shell Centre / MARS), background page
  8. Mathematics Assessment Project, Classroom Challenges lessons listing
  9. Malcolm Swan, "A Designer Speaks," Educational Designer, Vol.1 Issue 1, Article 3
  10. Fawn Nguyen, Visual Patterns
  11. Wikipedia, "Kakuro"
  12. Wikipedia, "KenKen"
  13. Wikipedia, "Fermi problem"
  14. Wikipedia, "Word problem (mathematics education)"
  15. Wikipedia, "The Nine Chapters on the Mathematical Art"
  16. Wikipedia, "Tangram"
  17. Wikipedia, "Vedic mathematics"
  18. Wikipedia, "Math circle"
  19. Wikipedia, "Automatic item generation"
  20. Math Strength, "Rich Tasks"

Open questions this document leaves for the owner

These are unanswered on purpose. They are listed, not resolved — turning them into a FAQ would mean inventing answers the document does not contain.

One of 51 research documents, 168,346 words in total, counted at build time from the files themselves. Read this document in the repository