Challenges and Grading
How a challenge is authored and how it will be graded: the Merlin Challenge
Config component, the grading approaches and strategies, question types,
answer keys, structured choice pools, and AI rubrics. Verified against POC
build 213 (20260910T123348).
Program (orchestrated) challenges have their own document — Programs and Workbench. This one covers everything reachable from the answer-based side.
Contents
- Missions and challenges
- The Merlin Challenge Config component
- Grading approach and strategy
- Question types and modifiers
- Answer keys
- Structured choice pools
- AI grading rubrics
- Naming and validation
- Variant versioning
Missions and challenges
A mission groups challenges. A challenge is a single graded question. Both are panel types in the lesson model, so they live in the same tree as content panels and are reordered the same way — with one structural rule: a challenge may only be created under a mission ancestor, and challenges may not nest inside one another.
The mission editor is host UI plus a MyST editor for the mission description. Each child challenge renders as a pill summarizing its grading strategy and question type, so a mission reads as a contents page for what it asks.
Every challenge carries an authored, minted-once challengeRef — the
join between the authoring tree and the published challenge shells. It is
minted when the challenge is created and never regenerated, because published
artifacts and workbench directories are addressed by it.
The Merlin Challenge Config component
A second published assembly (Merlin Challenge Config,
component-assemblies/challenge-config/), resolved from the same
published-components catalog as the editor and subject to the same version
pinning and local-override rules. It renders two faces:
- Instructor face — the grading rubric section and the Options pane.
- Student face — the response controls: a single-line input or a textarea, a dealt set of choices, Submit answer, Re-deal ↻, and a result area.
Which face shows follows the shell’s view mode. In instructor view the Options pane sits above the question, a view-scoped ordering chosen so the author sees how a question will be graded while writing it.
Grading approach and strategy
Two levels of choice, in two different places.
The host owns the outer choice, under the heading Grading approach:
| Label | evaluationStrategy |
|---|---|
| Answer-based | sync, aiSync or aiAsync |
| Program · orchestrated | jobAsync |
The component owns the inner choice for answer-based challenges, a segmented control labelled Grading:
| Label | evaluationStrategy |
What grades it |
|---|---|---|
| Deterministic | sync |
Exact comparison against the answer key, subject to the modifiers |
| AI · instant | aiSync |
A model grades against the rubric while the student waits |
| AI · deferred | aiAsync |
A model grades against the rubric out of band |
Switching to Program · orchestrated sets jobAsync and seeds the program
fields — a default runtime, a generator containing run.sh, and a submission
file name — without discarding the answer-based fields, so switching back
does not lose authored work. Switching away resets the strategy to sync.
Question types and modifiers
Question type applies to Deterministic grading (the Question type select disappears for AI strategies, which grade prose against a rubric):
| Label | evaluationMode |
|---|---|
| Short answer / fill-in-the-blank (text match) | textEquivalence |
| Multiple choice — one correct | singleSelect |
| Multiple choice — multiple correct | multipleSelect |
| Math expression (symbolic) | casEquivalence |
Two modifiers, both on by default, and both meaningless for symbolic matching — where the pane says so explicitly rather than showing dead controls:
| Label | Field |
|---|---|
| Ignore case | modifiers.caseInsensitive |
| Ignore extra whitespace | modifiers.ignoreInsignificantWhitespace |
Locale is en_US or es_MX. AI strategies additionally expose Max score
(default 10) and a Prompt bundle resource URN, defaulting to
urn:ayode:resource:/codermerlin.academy/system/ai/grade-short-freeform@1.
Answer keys
Accepted responses are full MyST source, not plain strings — so an answer key can carry maths, code or emphasis.
At rest each answer renders as its MyST output, using the same CDN renderer stack as the editor and degrading to plain text if that fails. Clicking a row swaps a shared inline editor into place; blur, Tab or Escape closes it. The open editor is scrolled into view and its section never flex-shrinks, so an answer row cannot be clipped by its container.
Structured choice pools
For the two select modes the author does not write an answer key but a choice pool — and the pool is authored file-first, ahead of the backend schema that will carry it.
Each choice is MyST with a stable choiceID and a correct flag. IDs are
minted once per choice and preserved, so recording which choice a student
picked survives the pool being reordered or extended.
Presentation controls how the pool is dealt to a student:
| Field | Meaning |
|---|---|
presentCount |
Deal m of n choices. null, or any value at or above the pool size, means deal all |
alwaysIncludeAccepted |
Every deal includes all correct choices (default on) |
shuffle |
Shuffle the dealt order (default on) |
presentCount is clamped rather than trusted: the floor is the number of
correct choices when alwaysIncludeAccepted is on, and 1 otherwise, so a
deal can never be constructed that omits the answer it will mark correct.
Empty choices are excluded from the pool entirely.
A live deal summary states in words what the student will get — for example “every deal: 1 correct + 3 of 7 distractors · order shuffled”, or “all 5 choices dealt · authored order”, or “add choices to preview the deal” when the pool is empty. Re-deal on the student face re-rolls it, so an author can watch the distribution rather than reason about it.
AI grading rubrics
For aiSync and aiAsync, a rubric grid replaces the answer key. Each
criterion has a stable criterionID, a label, a maximum score, and grading
guidance — the prose the model is asked to apply. The section footer sums the
criteria and shows the total beside + Add criterion, so a rubric that no
longer adds up to the challenge’s max score is visible while it is being
written.
Naming and validation
A new challenge is named from the nearest preceding content panel’s title —
<panel title> - Challenge <n>, with n advanced past any name already
taken — so a challenge arrives with a meaningful, unique name rather than a
placeholder.
Names are validated during authoring and the failures block publish,
because duplicate names in one generated-challenges request are a
server-side 400:
| Problem | Detected as |
|---|---|
| Duplicate | Two challenges with the same name, compared case-insensitively |
| Unnamed | Empty, or still literally “Untitled” or “New challenge” |
Both surface as authoring warnings first, so the author fixes them while writing rather than at the publish gate.
Variant versioning
The Options pane states the rule in the instructor’s own terms:
Saving creates a new challenge-variant version; earlier versions are kept and superseded, never edited in place.
That is a platform guarantee, not a UI nicety — a question a student has already answered is never rewritten underneath them. Challenge content is normalized and hashed (SHA-256 over canonical JSON) for dirty detection, so a save that changes nothing mints no new version; see Formats and Contracts.