ARFC-1010: Challenge Evidence and Scoring
Status
| Draft | Draft Date: 2026-08-08 | Last Call Date: — | Publication Date: — | Version: 0.1 |
Abstract
Students demonstrate what they have learned by attempting challenges — the smallest discrete pieces of work the platform delivers. Attempts are scored, one attempt per challenge is selected to count, and those scores roll up into grades. All of this is in production today and none of it is specified: an institution integrating against the platform must infer the rules by observing behavior.
This ARFC specifies that behavior without changing it, and then adds four properties that make a score interpretable as evidence of learning rather than merely a number: what kind of question was asked, how long the student worked on it, how much help the surrounding context gave, and whether the student had met that question before.
The distinction matters because the same score can mean opposite things. Ninety-five percent on a familiar problem sitting beside a worked example, with unlimited retries, is not the same accomplishment as ninety-five percent on an unfamiliar problem under exam conditions — yet today the platform records both identically. This ARFC makes the difference visible. It deliberately stops short of saying what the difference is worth; that is ARFC-1012.
Table of Contents
- Introduction
- Motivation
- Specification
- 3.1 Core Principles
- 3.2 Primary Objects
- 3.3 How a Score Is Formed
- 3.4 Which Attempt Counts
- 3.5 Required and Optional Work
- 3.6 Challenge Kind
- 3.7 Effort Time
- 3.8 Contextual Assistance
- 3.9 Novelty
- 3.10 Attempt Lifecycle
- 3.11 API Surface
- 3.12 Authorization Model
- 3.13 Absent, Unavailable, and Failed
- Version 1 Scope
- Security Considerations
- Backward Compatibility
- References
1. Introduction
A challenge is one question or task a student performs — a multiple-choice question, a fill-in-the-blank, a short written response, a program to write. Challenges are grouped into missions, missions into lessons, and a lesson is delivered to a class as an assignment with dates attached.
When a student submits work against a challenge, that submission is an attempt. Attempts are checked — automatically, by an AI reviewer, by a background job, or by a teacher — and a check produces a result: a score, on whatever scale that challenge uses, with optional feedback.
Everything above already works. What this ARFC addresses is that the rules governing it have never been written down, and that a result on its own does not record the conditions under which it was earned. Two students can hold identical results that represent entirely different levels of accomplishment, and nothing in the record distinguishes them.
This document is the first of a series. It specifies evidence. ARFC-1011 specifies what that evidence demonstrates about a learning objective. ARFC-1012 specifies how it becomes a reported grade.
2. Motivation
Current State
The following is live in both development and production, and is specified nowhere:
- A student may attempt a challenge more than once. Every attempt is retained.
- Each attempt may produce a result carrying a score, and the platform records what produced that result — an automated check, an AI reviewer, a background job, or a named teacher.
- A challenge’s score is on that challenge’s own scale. A right-or-wrong check scores out of one; a rubric-assessed response scores out of whatever the rubric defines.
- When a challenge is placed into an assignment, the instructor sets how much it is worth in that assignment, and chooses whether the student’s first, best, or last attempt is the one that counts.
- A placement is marked required or optional, and totals for the two are reported separately rather than blended.
An institution building against this must reverse-engineer all of it. Two consequences follow. Integrators guess at semantics that happen to be subtle — particularly how scores on differing scales are made comparable. And the platform has no agreed statement of intent to reconcile its own implementations against.
Missing Capabilities
Nothing records how long a student worked. The platform knows when an attempt was submitted and when checking finished, but that interval measures the grader, not the student. A student who answers in ten seconds and one who labours for twenty minutes are indistinguishable, so fluency cannot be observed at all.
Nothing records how much help the surroundings gave. A challenge placed immediately after an explanation, beside a nearly identical worked example, with unlimited retries, produces a result identical in every recorded respect to the same challenge in a closed end-of-unit assessment.
Nothing distinguishes a first encounter from a repeat. A student meeting a question for the third time may be demonstrating competence, or may be reproducing a remembered answer. The platform cannot tell, so anything built on top of it cannot either.
A challenge’s kind is not a dependable property. Whether a challenge asks a student to select an option or to explain their reasoning is not recorded in a form that reports and policies can rely on. A student who has only ever selected from lists has never explained anything — and no part of the platform can currently notice.
Design Goals
- Specify current behavior exactly, and change none of it. Everything in §3.3 through §3.5 describes what the platform already does.
- Make a score interpretable. The conditions under which a result was earned travel with the result, so that later interpretation is possible without re-deriving them.
- Ask instructors for nothing the platform can infer. Where a property can be determined from what the platform already knows, it is determined, not requested.
- Separate recording from judging. This ARFC records what happened. It assigns no meaning, sets no thresholds, and defines no levels.
- Be purely additive. No existing behavior changes, no existing integration breaks, and no historical result is reinterpreted.
3. Specification
3.1 Core Principles
Evidence is recorded at the smallest unit of work. A challenge is where a score is earned and where its conditions are recorded. Everything larger — a mission, a lesson, a course — is a summary of the challenges beneath it, never a place where evidence is entered directly. This mirrors how standards coverage already works on the platform, and it is what makes a gap findable rather than merely asserted.
One score, many meanings. A number alone is not evidence. The conditions of its earning are part of the record, not metadata about it.
The platform infers what it can; authors state only what only they know. Whether a student has met a question before is something the platform already knows and must never ask. How much a lesson’s surrounding material gives away is a judgement only the author can make, so the platform proposes and the author decides.
Evidence is never overwritten. Every attempt is retained. Selecting which attempt counts is an interpretation applied over a complete history, not a deletion of the rest.
Recording is separate from interpreting. The same recorded evidence supports a traditional percentage and a mastery judgement without being stored twice or reconciled between them.
3.2 Primary Objects
Challenge — one question or task. Belongs to a mission. Independent of any class, student, or date.
Challenge kind — what form the challenge takes, and therefore what it can and cannot demonstrate. See §3.6.
Attempt — one submission by one student against one challenge in one assignment. Numbered in sequence and retained permanently.
Result — the outcome of checking an attempt: a score on the challenge’s own scale, what produced it, when, and optional feedback. At most one result per attempt.
Placement — a challenge as it appears in one particular assignment. The placement, not the challenge, carries how much the work is worth, whether it is required, which attempt counts, and how much contextual assistance surrounds it. The same challenge placed in two assignments has two placements and may behave completely differently in each.
Evidence conditions — three properties describing the circumstances of an attempt:
| Condition | What it describes | Where it comes from |
|---|---|---|
| Effort time | how long the student worked | recorded as the student works |
| Contextual assistance | how much the surroundings helped | set by the author on the placement |
| Novelty | whether the student had met this challenge before | inferred from the student’s own history |
That the three have different origins is deliberate and load-bearing: one is observed, one is declared, one is derived. None can be substituted for another.
flowchart LR
C[Challenge<br/>a question or task] -->|placed into an assignment| P[Placement<br/>worth · required · which attempt counts<br/>· contextual assistance]
P -->|student submits| A[Attempt<br/>numbered · retained · effort time]
A -->|checked| R[Result<br/>score on the challenge's own scale]
A -.->|compared with the student's own history| N[Novelty]
R --> E[Evidence<br/>score plus its conditions]
N --> E
3.3 How a Score Is Formed
A placement’s contribution to a grade is the proportion of the challenge the student got right, multiplied by what that placement is worth.
The proportion step exists because challenges are scored on their own scales. A right-or-wrong check scores out of one; a rubric-assessed essay might score out of forty. Multiplying raw scores by weights would make the essay silently dominate. Taking the proportion first means three out of four and thirty out of forty contribute identically — which is the property that lets an instructor mix question types in one assignment without distorting the result.
What a placement is worth is set per placement. The consequence is that the same challenge can be ungraded practice in one class and a major grade in another, without being edited — and a challenge therefore carries no intrinsic value of its own.
A mission’s total possible score is the sum of what its placements are worth. A student’s earned score is the sum of their contributions.
3.4 Which Attempt Counts
Students normally attempt a challenge several times. Exactly one attempt is selected to count, per the instructor’s choice:
| Choice | Behavior | Typical use |
|---|---|---|
| First | the earliest attempt counts | a diagnostic, where later practice must not inflate the reading |
| Best | the highest-scoring attempt counts | practice, where the goal is eventual success |
| Last | the most recent attempt counts | revision, where the current state of understanding is the point |
Three properties are normative:
- The choice is per placement, and explicit. There is no platform default. Composing an assignment requires stating the choice for each challenge, because a silent default would quietly decide how students are assessed.
- Selection is deterministic. When two attempts tie under the chosen rule, exactly one is selected, and the same one is selected every time the calculation runs. A tie must never produce two counting attempts or a result that varies between views of the same data.
- Nothing is discarded. Unselected attempts remain part of the record and remain visible.
3.5 Required and Optional Work
Every placement is either required or optional, and the two are reported as separate totals that are never merged.
An instructor sees earned and possible scores, and completed and total counts, for required work and for optional work independently. A blended figure would hide the distinction that matters most — whether a student has done what was asked, or has done a lot of enrichment while leaving requirements unmet.
3.6 Challenge Kind
A challenge’s kind is a first-class property that reports and policies may rely on:
| Kind | The student… |
|---|---|
| Single select | chooses one option from several |
| Multiple select | chooses every applicable option |
| Fill in the blank | supplies a short exact answer |
| Expression | supplies a mathematical expression judged for equivalence rather than literal match |
| Short answer | writes a brief response in their own words |
| Essay | writes an extended response |
Kind determines what a challenge is capable of demonstrating. Only the last two can show that a student can explain or justify anything; the first four cannot, however many of them a student answers correctly. A platform that cannot see the difference will eventually report an understanding it never observed.
Kind also sets the expectation against which effort time is interpreted (§3.7).
Kind is a property of the challenge, not of its placement — it does not change with context.
3.7 Effort Time
Effort time is the time attributable to the student’s own work on an attempt.
It is explicitly not the time the platform took to check the work. That interval is a property of the grading system and says nothing about the student.
Recording must survive realistic behavior. Students open work and abandon it, resume days later, work across several sittings, and leave a screen open while doing something else. A single measurement from first opening to final submission would overstate effort badly and consistently.
Effort time is meaningless in the absolute and is interpreted against an expected duration for the challenge’s kind. Forty seconds is fluent for a single-select question and implausible for an essay. Expected durations are a platform-level expectation per kind, adjustable by an institution.
Effort time is recorded per attempt, so a student growing faster across attempts is visible as such.
3.8 Contextual Assistance
Contextual assistance describes how much a challenge’s surroundings help a student answer it.
A challenge inserted mid-lesson, immediately after an explanation and beside a nearly identical worked example, is heavily assisted: much of the method is on the page. The same challenge in a closed end-of-unit assessment is barely assisted at all. Assistance is a property of placement, not of the challenge — which is precisely why it belongs on the placement.
Four tiers, ordered from most to least assistance:
| Tier | The student meets the challenge… |
|---|---|
| Scaffolded | with explanation and a near-identical worked example adjacent; the method is essentially shown |
| Practiced | in familiar framing, within the same lesson or unit, with retries permitted |
| Standard | in an expected context, without adjacent instruction to lean on |
| Unaided | in unfamiliar framing, away from instruction, with a single opportunity |
The tier is set by the author, with a value proposed by the platform. The platform proposes from what it can see about the placement: whether retries are permitted, what the placement is worth relative to its siblings, where it sits relative to the surrounding instruction, and whether this is the lesson that introduced the material. The author confirms or changes it. A proposal is a labour saving, never a decision — a tier that shapes reported attainment is not something to infer silently.
Assistance describes the placement as authored, not any individual student’s experience of it. Live human help — a teacher at the desk, a neighbour explaining — is outside this version’s scope (§4).
ARFC-1011 refines assistance per learning objective, since one challenge may be heavily assisted with respect to one objective and barely assisted with respect to another.
3.9 Novelty
Novelty describes whether a student had already met this challenge before this attempt.
It is inferred entirely from what the student has already done and requires no author or instructor action whatsoever. The platform already knows every challenge every student has attempted.
Three tiers:
| Tier | The student… |
|---|---|
| Repeat | has met this challenge in this form before |
| Reformed | has met this challenge before, but in a different form |
| Novel | has never met this challenge before |
The middle tier reflects an existing capability: the platform can present different forms of the same question to different students. A student meeting a different form is doing neither a pure repeat nor a fresh encounter, and collapsing that into a binary would discard a real distinction the platform already creates.
Novelty is a property of the student’s history, not of the placement. Two students meeting the same challenge in the same assignment may hold different novelty tiers, and a student’s tier for a given challenge can only decrease over time as they meet it again.
Novelty and assistance are independent and must not be conflated. A student may meet a familiar question under exam conditions — low assistance, but a repeat — and that combination is emphatically not the same accomplishment as meeting an unfamiliar question under the same conditions.
flowchart TB
subgraph observed [Observed as the student works]
T[Effort time]
end
subgraph declared [Declared by the author, platform proposes]
A[Contextual assistance<br/>scaffolded → practiced → standard → unaided]
end
subgraph derived [Derived from the student's own history]
N[Novelty<br/>repeat → reformed → novel]
end
T --> EV[Evidence for one attempt]
A --> EV
N --> EV
S[Score, as a proportion] --> EV
K[Challenge kind] --> EV
3.10 Attempt Lifecycle
stateDiagram-v2
[*] --> Submitted: student submits
Submitted --> Running: checking begins
Running --> Completed: checking succeeds
Running --> Failed: checking cannot complete
Completed --> [*]
Failed --> [*]
note right of Running
A freshly submitted attempt may
briefly show no score. This is
checking in progress, not a zero.
end note
A failed attempt is one the platform could not check. It is not a score of zero and must never be presented as one; the student’s work may be perfectly correct. Failure is surfaced as failure (§3.13).
3.11 API Surface
Existing assignment-composition and result-reading operations gain optional fields rather than being replaced. No existing operation changes shape for a client that ignores the additions.
| Method | Path | operationId | Description |
|---|---|---|---|
| PATCH | /v2/assignments/{assignment-eid}/challenges/{placement-eid} |
updateAssignmentChallengeAssistanceV2 |
Set a placement’s contextual assistance tier |
| GET | /v2/assignments/{assignment-eid}/challenges/{placement-eid}/assistance-proposal |
getAssignmentChallengeAssistanceProposalV2 |
Retrieve the platform’s proposed tier, with the reasons behind it |
| GET | /v2/challenges/{challenge-eid}/evidence-conditions |
getChallengeEvidenceConditionsV2 |
Read a challenge’s kind and expected duration |
| GET | /v2/sections/{section-eid}/students/{student-eid}/attempts |
listStudentAttemptEvidenceV2 |
List a student’s attempts with score, effort time, assistance, and novelty |
The assistance proposal returns its reasoning, not merely a value. An author asked to confirm a judgement is owed the basis for it.
3.12 Authorization Model
Setting contextual assistance requires the same authority as composing the assignment: an instructor may set it for their own sections, and sees it read-only elsewhere. Because the tier shapes reported attainment, an institution may restrict it further — reserving it to a curriculum lead — in which case instructors see the tier but cannot change it.
Effort time and novelty are platform-recorded and cannot be set by anyone. There is no endpoint to write them. This is deliberate: they are the two conditions that would be most tempting to adjust in a student’s favour, and making them unwritable removes the question.
Reading evidence follows existing scope. A student may read their own evidence and no one else’s. An instructor may read evidence for students in sections they are assigned to. A request outside that scope is refused without revealing whether the student or section exists.
Novelty must not leak. A student’s novelty tier is derived from their own history only. No response may allow a reader to infer what other students have attempted.
3.13 Absent, Unavailable, and Failed
Four states are distinct and must never be collapsed into one another:
| State | Meaning | Must not be shown as |
|---|---|---|
| No attempt | the student has not attempted the work | a score of zero |
| Checking | an attempt is being checked | a score of zero |
| Failed | the platform could not check the attempt | a score of zero |
| Not recorded | a condition was never captured for this attempt | the lowest tier, or a zero |
The last deserves emphasis. Historical attempts carry no effort time and no assistance tier. Presenting an unrecorded assistance tier as scaffolded, or an unrecorded effort time as zero, would silently misrepresent past work. Absent means absent.
An attempt missing one condition still supplies the others. A result with no effort time is fully usable as evidence of accuracy; only fluency is unavailable for it.
4. Version 1 Scope
Included
- Specification of existing scoring: proportion-and-weight, attempt selection, required/optional separation
- Deterministic tie behavior in attempt selection
- Challenge kind as a first-class property
- Effort time, tolerant of abandonment, resumption, and multiple sittings
- Expected duration per challenge kind, adjustable by an institution
- Contextual assistance: four tiers, author-set, platform-proposed with reasons
- Novelty: three tiers, wholly inferred
- Distinct handling of no-attempt, checking, failed, and not-recorded
Excluded
- Mastery levels and what evidence is worth — ARFC-1012
- Learning objectives, alignment, and per-objective assistance refinement — ARFC-1011
- Grading policies, scales, and reported grades — ARFC-1012
- Live human assistance, whether from a teacher or a peer
- Whether a student independently recognised and corrected their own error
- Peer instruction as evidence
- Retroactive effort time for historical attempts, which cannot be reconstructed
5. Security Considerations
Effort time is behavioral data about a minor. How long a child took to answer a question is more revealing than the answer, and invites inference about attention and difficulty that the platform does not warrant. It carries the same access scope as scores, is never exposed to peers, and is not aggregated into anything comparative across students in this version.
Contextual assistance is a grading-integrity control, not a convenience setting. Lowering a placement’s tier raises the apparent accomplishment of every student who attempted it. Authority to set it must be governed accordingly, and every change must be attributable to the person who made it.
Effort time and novelty are unwritable by design. Making the two most sensitive conditions platform-derived removes any question of whether a recorded condition was adjusted after the fact.
Novelty must not become a channel for other students’ activity. Because it is derived from history, an implementation that resolves it too broadly could allow one student’s response to reveal another’s work. Novelty is derived from the requesting student’s history alone.
Evidence conditions are part of the academic record. They are subject to the same retention, audit, and disclosure obligations as scores, including a parent’s or student’s right to understand how a grade was reached.
6. Backward Compatibility
Nothing existing changes. Scoring, attempt selection, and required/optional reporting are specified exactly as they behave. No score is recalculated, no grade moves, and no client sees a different value for anything it reads today.
Additions are optional. A client that ignores the new properties continues to work unchanged. New fields are additive to existing responses; no field changes meaning or type.
Historical work has no conditions, and must not be given defaults. Existing attempts carry no effort time and their placements carry no assistance tier. These are not recorded, not zero and not the lowest tier. Defaulting historical placements to scaffolded would cap the apparent attainment of every student in the platform’s history — a silent, large, and irreversible misrepresentation. Absent conditions are reported absent, and anything consuming them treats the corresponding evidence as unavailable rather than negative.
The two new conditions differ in whether they can be backfilled. Novelty is derivable retroactively, because the attempt history it rests on already exists — historical attempts can be resolved without new information. Effort time cannot: it was never observed and cannot be reconstructed. Fluency evidence therefore begins accumulating only from the point recording starts, and any measure depending on it must tolerate its absence for existing students indefinitely.
Existing tie behavior is being made consistent, not changed in intent. Two implementations of attempt selection currently disagree on fully tied attempts. §3.4 states the intended behavior so reconciliation has a specification to converge on rather than a choice to make.
7. References
Related ARFCs
- ARFC-1011: Learning Objectives and Challenge Alignment — what evidence demonstrates about an objective
- ARFC-1012: Grading Policy and Rollup — what evidence is worth, and how it is reported
- ARFC-1004: Teams, Leagues, Seasons, and Competitive Gaming Schema — establishes the principle this ARFC follows, that aggregates are derived from an authoritative evidence record and recomputed when it changes, rather than being independently maintained
- ARFC-1007: Contextual Chat System — structural model for this document
Related Issues
codermerlin.academy-backend#3700— capture per-attempt time-on-task so challenge fluency can be measuredcodermerlin.academy-backend#3547— the deployed mission grade block, including required/optional separationcodermerlin.academy-backend#3546— the deployed attempt-selection behaviorcodermerlin.academy-backend#3544— reconciliation of the two attempt-selection implementationscodermerlin.academy-backend#2453— incomplete job-based grading, which bears on extended written responsescodermerlin.academy-backend#3253— objective progress and coverage, blocked pending the definitions in ARFC-1011 and ARFC-1012
Platform Documentation
Author
ĀYŌDÈ Development Team Codermerlin Academy Architecture