Challenges and Grading

How a challenge is authored and how it will be graded: the Merlin Challenge Config component, the grading approaches and strategies, question types, answer keys, structured choice pools, and AI rubrics. Verified against POC build 213 (20260910T123348).

Program (orchestrated) challenges have their own document — Programs and Workbench. This one covers everything reachable from the answer-based side.

Contents

Missions and challenges

A mission groups challenges. A challenge is a single graded question. Both are panel types in the lesson model, so they live in the same tree as content panels and are reordered the same way — with one structural rule: a challenge may only be created under a mission ancestor, and challenges may not nest inside one another.

The mission editor is host UI plus a MyST editor for the mission description. Each child challenge renders as a pill summarizing its grading strategy and question type, so a mission reads as a contents page for what it asks.

Every challenge carries an authored, minted-once challengeRef — the join between the authoring tree and the published challenge shells. It is minted when the challenge is created and never regenerated, because published artifacts and workbench directories are addressed by it.

The Merlin Challenge Config component

A second published assembly (Merlin Challenge Config, component-assemblies/challenge-config/), resolved from the same published-components catalog as the editor and subject to the same version pinning and local-override rules. It renders two faces:

Which face shows follows the shell’s view mode. In instructor view the Options pane sits above the question, a view-scoped ordering chosen so the author sees how a question will be graded while writing it.

Grading approach and strategy

Two levels of choice, in two different places.

The host owns the outer choice, under the heading Grading approach:

Label evaluationStrategy
Answer-based sync, aiSync or aiAsync
Program · orchestrated jobAsync

The component owns the inner choice for answer-based challenges, a segmented control labelled Grading:

Label evaluationStrategy What grades it
Deterministic sync Exact comparison against the answer key, subject to the modifiers
AI · instant aiSync A model grades against the rubric while the student waits
AI · deferred aiAsync A model grades against the rubric out of band

Switching to Program · orchestrated sets jobAsync and seeds the program fields — a default runtime, a generator containing run.sh, and a submission file name — without discarding the answer-based fields, so switching back does not lose authored work. Switching away resets the strategy to sync.

Question types and modifiers

Question type applies to Deterministic grading (the Question type select disappears for AI strategies, which grade prose against a rubric):

Label evaluationMode
Short answer / fill-in-the-blank (text match) textEquivalence
Multiple choice — one correct singleSelect
Multiple choice — multiple correct multipleSelect
Math expression (symbolic) casEquivalence

Two modifiers, both on by default, and both meaningless for symbolic matching — where the pane says so explicitly rather than showing dead controls:

Label Field
Ignore case modifiers.caseInsensitive
Ignore extra whitespace modifiers.ignoreInsignificantWhitespace

Locale is en_US or es_MX. AI strategies additionally expose Max score (default 10) and a Prompt bundle resource URN, defaulting to urn:ayode:resource:/codermerlin.academy/system/ai/grade-short-freeform@1.

Answer keys

Accepted responses are full MyST source, not plain strings — so an answer key can carry maths, code or emphasis.

At rest each answer renders as its MyST output, using the same CDN renderer stack as the editor and degrading to plain text if that fails. Clicking a row swaps a shared inline editor into place; blur, Tab or Escape closes it. The open editor is scrolled into view and its section never flex-shrinks, so an answer row cannot be clipped by its container.

Structured choice pools

For the two select modes the author does not write an answer key but a choice pool — and the pool is authored file-first, ahead of the backend schema that will carry it.

Each choice is MyST with a stable choiceID and a correct flag. IDs are minted once per choice and preserved, so recording which choice a student picked survives the pool being reordered or extended.

Presentation controls how the pool is dealt to a student:

Field Meaning
presentCount Deal m of n choices. null, or any value at or above the pool size, means deal all
alwaysIncludeAccepted Every deal includes all correct choices (default on)
shuffle Shuffle the dealt order (default on)

presentCount is clamped rather than trusted: the floor is the number of correct choices when alwaysIncludeAccepted is on, and 1 otherwise, so a deal can never be constructed that omits the answer it will mark correct. Empty choices are excluded from the pool entirely.

A live deal summary states in words what the student will get — for example “every deal: 1 correct + 3 of 7 distractors · order shuffled”, or “all 5 choices dealt · authored order”, or “add choices to preview the deal” when the pool is empty. Re-deal on the student face re-rolls it, so an author can watch the distribution rather than reason about it.

AI grading rubrics

For aiSync and aiAsync, a rubric grid replaces the answer key. Each criterion has a stable criterionID, a label, a maximum score, and grading guidance — the prose the model is asked to apply. The section footer sums the criteria and shows the total beside + Add criterion, so a rubric that no longer adds up to the challenge’s max score is visible while it is being written.

Naming and validation

A new challenge is named from the nearest preceding content panel’s title — <panel title> - Challenge <n>, with n advanced past any name already taken — so a challenge arrives with a meaningful, unique name rather than a placeholder.

Names are validated during authoring and the failures block publish, because duplicate names in one generated-challenges request are a server-side 400:

Problem Detected as
Duplicate Two challenges with the same name, compared case-insensitively
Unnamed Empty, or still literally “Untitled” or “New challenge”

Both surface as authoring warnings first, so the author fixes them while writing rather than at the publish gate.

Variant versioning

The Options pane states the rule in the instructor’s own terms:

Saving creates a new challenge-variant version; earlier versions are kept and superseded, never edited in place.

That is a platform guarantee, not a UI nicety — a question a student has already answered is never rewritten underneath them. Challenge content is normalized and hashed (SHA-256 over canonical JSON) for dirty detection, so a save that changes nothing mints no new version; see Formats and Contracts.