ARFC-1010: Challenge Evidence and Scoring

Status

Draft Draft Date: 2026-08-08 Last Call Date: — Publication Date: — Version: 0.1

Abstract

Students demonstrate what they have learned by attempting challenges — the smallest discrete pieces of work the platform delivers. Attempts are scored, one attempt per challenge is selected to count, and those scores roll up into grades. All of this is in production today and none of it is specified: an institution integrating against the platform must infer the rules by observing behavior.

This ARFC specifies that behavior without changing it, and then adds four properties that make a score interpretable as evidence of learning rather than merely a number: what kind of question was asked, how long the student worked on it, how much help the surrounding context gave, and whether the student had met that question before.

The distinction matters because the same score can mean opposite things. Ninety-five percent on a familiar problem sitting beside a worked example, with unlimited retries, is not the same accomplishment as ninety-five percent on an unfamiliar problem under exam conditions — yet today the platform records both identically. This ARFC makes the difference visible. It deliberately stops short of saying what the difference is worth; that is ARFC-1012.


Table of Contents

  1. Introduction
  2. Motivation
  3. Specification
  4. Version 1 Scope
  5. Security Considerations
  6. Backward Compatibility
  7. References

1. Introduction

A challenge is one question or task a student performs — a multiple-choice question, a fill-in-the-blank, a short written response, a program to write. Challenges are grouped into missions, missions into lessons, and a lesson is delivered to a class as an assignment with dates attached.

When a student submits work against a challenge, that submission is an attempt. Attempts are checked — automatically, by an AI reviewer, by a background job, or by a teacher — and a check produces a result: a score, on whatever scale that challenge uses, with optional feedback.

Everything above already works. What this ARFC addresses is that the rules governing it have never been written down, and that a result on its own does not record the conditions under which it was earned. Two students can hold identical results that represent entirely different levels of accomplishment, and nothing in the record distinguishes them.

This document is the first of a series. It specifies evidence. ARFC-1011 specifies what that evidence demonstrates about a learning objective. ARFC-1012 specifies how it becomes a reported grade.


2. Motivation

Current State

The following is live in both development and production, and is specified nowhere:

An institution building against this must reverse-engineer all of it. Two consequences follow. Integrators guess at semantics that happen to be subtle — particularly how scores on differing scales are made comparable. And the platform has no agreed statement of intent to reconcile its own implementations against.

Missing Capabilities

Nothing records how long a student worked. The platform knows when an attempt was submitted and when checking finished, but that interval measures the grader, not the student. A student who answers in ten seconds and one who labours for twenty minutes are indistinguishable, so fluency cannot be observed at all.

Nothing records how much help the surroundings gave. A challenge placed immediately after an explanation, beside a nearly identical worked example, with unlimited retries, produces a result identical in every recorded respect to the same challenge in a closed end-of-unit assessment.

Nothing distinguishes a first encounter from a repeat. A student meeting a question for the third time may be demonstrating competence, or may be reproducing a remembered answer. The platform cannot tell, so anything built on top of it cannot either.

A challenge’s kind is not a dependable property. Whether a challenge asks a student to select an option or to explain their reasoning is not recorded in a form that reports and policies can rely on. A student who has only ever selected from lists has never explained anything — and no part of the platform can currently notice.

Design Goals

  1. Specify current behavior exactly, and change none of it. Everything in §3.3 through §3.5 describes what the platform already does.
  2. Make a score interpretable. The conditions under which a result was earned travel with the result, so that later interpretation is possible without re-deriving them.
  3. Ask instructors for nothing the platform can infer. Where a property can be determined from what the platform already knows, it is determined, not requested.
  4. Separate recording from judging. This ARFC records what happened. It assigns no meaning, sets no thresholds, and defines no levels.
  5. Be purely additive. No existing behavior changes, no existing integration breaks, and no historical result is reinterpreted.

3. Specification

3.1 Core Principles

Evidence is recorded at the smallest unit of work. A challenge is where a score is earned and where its conditions are recorded. Everything larger — a mission, a lesson, a course — is a summary of the challenges beneath it, never a place where evidence is entered directly. This mirrors how standards coverage already works on the platform, and it is what makes a gap findable rather than merely asserted.

One score, many meanings. A number alone is not evidence. The conditions of its earning are part of the record, not metadata about it.

The platform infers what it can; authors state only what only they know. Whether a student has met a question before is something the platform already knows and must never ask. How much a lesson’s surrounding material gives away is a judgement only the author can make, so the platform proposes and the author decides.

Evidence is never overwritten. Every attempt is retained. Selecting which attempt counts is an interpretation applied over a complete history, not a deletion of the rest.

Recording is separate from interpreting. The same recorded evidence supports a traditional percentage and a mastery judgement without being stored twice or reconciled between them.

3.2 Primary Objects

Challenge — one question or task. Belongs to a mission. Independent of any class, student, or date.

Challenge kind — what form the challenge takes, and therefore what it can and cannot demonstrate. See §3.6.

Attempt — one submission by one student against one challenge in one assignment. Numbered in sequence and retained permanently.

Result — the outcome of checking an attempt: a score on the challenge’s own scale, what produced it, when, and optional feedback. At most one result per attempt.

Placement — a challenge as it appears in one particular assignment. The placement, not the challenge, carries how much the work is worth, whether it is required, which attempt counts, and how much contextual assistance surrounds it. The same challenge placed in two assignments has two placements and may behave completely differently in each.

Evidence conditions — three properties describing the circumstances of an attempt:

Condition What it describes Where it comes from
Effort time how long the student worked recorded as the student works
Contextual assistance how much the surroundings helped set by the author on the placement
Novelty whether the student had met this challenge before inferred from the student’s own history

That the three have different origins is deliberate and load-bearing: one is observed, one is declared, one is derived. None can be substituted for another.

flowchart LR
    C[Challenge<br/>a question or task] -->|placed into an assignment| P[Placement<br/>worth · required · which attempt counts<br/>· contextual assistance]
    P -->|student submits| A[Attempt<br/>numbered · retained · effort time]
    A -->|checked| R[Result<br/>score on the challenge's own scale]
    A -.->|compared with the student's own history| N[Novelty]
    R --> E[Evidence<br/>score plus its conditions]
    N --> E

3.3 How a Score Is Formed

A placement’s contribution to a grade is the proportion of the challenge the student got right, multiplied by what that placement is worth.

The proportion step exists because challenges are scored on their own scales. A right-or-wrong check scores out of one; a rubric-assessed essay might score out of forty. Multiplying raw scores by weights would make the essay silently dominate. Taking the proportion first means three out of four and thirty out of forty contribute identically — which is the property that lets an instructor mix question types in one assignment without distorting the result.

What a placement is worth is set per placement. The consequence is that the same challenge can be ungraded practice in one class and a major grade in another, without being edited — and a challenge therefore carries no intrinsic value of its own.

A mission’s total possible score is the sum of what its placements are worth. A student’s earned score is the sum of their contributions.

3.4 Which Attempt Counts

Students normally attempt a challenge several times. Exactly one attempt is selected to count, per the instructor’s choice:

Choice Behavior Typical use
First the earliest attempt counts a diagnostic, where later practice must not inflate the reading
Best the highest-scoring attempt counts practice, where the goal is eventual success
Last the most recent attempt counts revision, where the current state of understanding is the point

Three properties are normative:

3.5 Required and Optional Work

Every placement is either required or optional, and the two are reported as separate totals that are never merged.

An instructor sees earned and possible scores, and completed and total counts, for required work and for optional work independently. A blended figure would hide the distinction that matters most — whether a student has done what was asked, or has done a lot of enrichment while leaving requirements unmet.

3.6 Challenge Kind

A challenge’s kind is a first-class property that reports and policies may rely on:

Kind The student…
Single select chooses one option from several
Multiple select chooses every applicable option
Fill in the blank supplies a short exact answer
Expression supplies a mathematical expression judged for equivalence rather than literal match
Short answer writes a brief response in their own words
Essay writes an extended response

Kind determines what a challenge is capable of demonstrating. Only the last two can show that a student can explain or justify anything; the first four cannot, however many of them a student answers correctly. A platform that cannot see the difference will eventually report an understanding it never observed.

Kind also sets the expectation against which effort time is interpreted (§3.7).

Kind is a property of the challenge, not of its placement — it does not change with context.

3.7 Effort Time

Effort time is the time attributable to the student’s own work on an attempt.

It is explicitly not the time the platform took to check the work. That interval is a property of the grading system and says nothing about the student.

Recording must survive realistic behavior. Students open work and abandon it, resume days later, work across several sittings, and leave a screen open while doing something else. A single measurement from first opening to final submission would overstate effort badly and consistently.

Effort time is meaningless in the absolute and is interpreted against an expected duration for the challenge’s kind. Forty seconds is fluent for a single-select question and implausible for an essay. Expected durations are a platform-level expectation per kind, adjustable by an institution.

Effort time is recorded per attempt, so a student growing faster across attempts is visible as such.

3.8 Contextual Assistance

Contextual assistance describes how much a challenge’s surroundings help a student answer it.

A challenge inserted mid-lesson, immediately after an explanation and beside a nearly identical worked example, is heavily assisted: much of the method is on the page. The same challenge in a closed end-of-unit assessment is barely assisted at all. Assistance is a property of placement, not of the challenge — which is precisely why it belongs on the placement.

Four tiers, ordered from most to least assistance:

Tier The student meets the challenge…
Scaffolded with explanation and a near-identical worked example adjacent; the method is essentially shown
Practiced in familiar framing, within the same lesson or unit, with retries permitted
Standard in an expected context, without adjacent instruction to lean on
Unaided in unfamiliar framing, away from instruction, with a single opportunity

The tier is set by the author, with a value proposed by the platform. The platform proposes from what it can see about the placement: whether retries are permitted, what the placement is worth relative to its siblings, where it sits relative to the surrounding instruction, and whether this is the lesson that introduced the material. The author confirms or changes it. A proposal is a labour saving, never a decision — a tier that shapes reported attainment is not something to infer silently.

Assistance describes the placement as authored, not any individual student’s experience of it. Live human help — a teacher at the desk, a neighbour explaining — is outside this version’s scope (§4).

ARFC-1011 refines assistance per learning objective, since one challenge may be heavily assisted with respect to one objective and barely assisted with respect to another.

3.9 Novelty

Novelty describes whether a student had already met this challenge before this attempt.

It is inferred entirely from what the student has already done and requires no author or instructor action whatsoever. The platform already knows every challenge every student has attempted.

Three tiers:

Tier The student…
Repeat has met this challenge in this form before
Reformed has met this challenge before, but in a different form
Novel has never met this challenge before

The middle tier reflects an existing capability: the platform can present different forms of the same question to different students. A student meeting a different form is doing neither a pure repeat nor a fresh encounter, and collapsing that into a binary would discard a real distinction the platform already creates.

Novelty is a property of the student’s history, not of the placement. Two students meeting the same challenge in the same assignment may hold different novelty tiers, and a student’s tier for a given challenge can only decrease over time as they meet it again.

Novelty and assistance are independent and must not be conflated. A student may meet a familiar question under exam conditions — low assistance, but a repeat — and that combination is emphatically not the same accomplishment as meeting an unfamiliar question under the same conditions.

flowchart TB
    subgraph observed [Observed as the student works]
        T[Effort time]
    end
    subgraph declared [Declared by the author, platform proposes]
        A[Contextual assistance<br/>scaffolded → practiced → standard → unaided]
    end
    subgraph derived [Derived from the student's own history]
        N[Novelty<br/>repeat → reformed → novel]
    end
    T --> EV[Evidence for one attempt]
    A --> EV
    N --> EV
    S[Score, as a proportion] --> EV
    K[Challenge kind] --> EV

3.10 Attempt Lifecycle

stateDiagram-v2
    [*] --> Submitted: student submits
    Submitted --> Running: checking begins
    Running --> Completed: checking succeeds
    Running --> Failed: checking cannot complete
    Completed --> [*]
    Failed --> [*]
    note right of Running
        A freshly submitted attempt may
        briefly show no score. This is
        checking in progress, not a zero.
    end note

A failed attempt is one the platform could not check. It is not a score of zero and must never be presented as one; the student’s work may be perfectly correct. Failure is surfaced as failure (§3.13).

3.11 API Surface

Existing assignment-composition and result-reading operations gain optional fields rather than being replaced. No existing operation changes shape for a client that ignores the additions.

Method Path operationId Description
PATCH /v2/assignments/{assignment-eid}/challenges/{placement-eid} updateAssignmentChallengeAssistanceV2 Set a placement’s contextual assistance tier
GET /v2/assignments/{assignment-eid}/challenges/{placement-eid}/assistance-proposal getAssignmentChallengeAssistanceProposalV2 Retrieve the platform’s proposed tier, with the reasons behind it
GET /v2/challenges/{challenge-eid}/evidence-conditions getChallengeEvidenceConditionsV2 Read a challenge’s kind and expected duration
GET /v2/sections/{section-eid}/students/{student-eid}/attempts listStudentAttemptEvidenceV2 List a student’s attempts with score, effort time, assistance, and novelty

The assistance proposal returns its reasoning, not merely a value. An author asked to confirm a judgement is owed the basis for it.

3.12 Authorization Model

Setting contextual assistance requires the same authority as composing the assignment: an instructor may set it for their own sections, and sees it read-only elsewhere. Because the tier shapes reported attainment, an institution may restrict it further — reserving it to a curriculum lead — in which case instructors see the tier but cannot change it.

Effort time and novelty are platform-recorded and cannot be set by anyone. There is no endpoint to write them. This is deliberate: they are the two conditions that would be most tempting to adjust in a student’s favour, and making them unwritable removes the question.

Reading evidence follows existing scope. A student may read their own evidence and no one else’s. An instructor may read evidence for students in sections they are assigned to. A request outside that scope is refused without revealing whether the student or section exists.

Novelty must not leak. A student’s novelty tier is derived from their own history only. No response may allow a reader to infer what other students have attempted.

3.13 Absent, Unavailable, and Failed

Four states are distinct and must never be collapsed into one another:

State Meaning Must not be shown as
No attempt the student has not attempted the work a score of zero
Checking an attempt is being checked a score of zero
Failed the platform could not check the attempt a score of zero
Not recorded a condition was never captured for this attempt the lowest tier, or a zero

The last deserves emphasis. Historical attempts carry no effort time and no assistance tier. Presenting an unrecorded assistance tier as scaffolded, or an unrecorded effort time as zero, would silently misrepresent past work. Absent means absent.

An attempt missing one condition still supplies the others. A result with no effort time is fully usable as evidence of accuracy; only fluency is unavailable for it.


4. Version 1 Scope

Included

Excluded


5. Security Considerations

Effort time is behavioral data about a minor. How long a child took to answer a question is more revealing than the answer, and invites inference about attention and difficulty that the platform does not warrant. It carries the same access scope as scores, is never exposed to peers, and is not aggregated into anything comparative across students in this version.

Contextual assistance is a grading-integrity control, not a convenience setting. Lowering a placement’s tier raises the apparent accomplishment of every student who attempted it. Authority to set it must be governed accordingly, and every change must be attributable to the person who made it.

Effort time and novelty are unwritable by design. Making the two most sensitive conditions platform-derived removes any question of whether a recorded condition was adjusted after the fact.

Novelty must not become a channel for other students’ activity. Because it is derived from history, an implementation that resolves it too broadly could allow one student’s response to reveal another’s work. Novelty is derived from the requesting student’s history alone.

Evidence conditions are part of the academic record. They are subject to the same retention, audit, and disclosure obligations as scores, including a parent’s or student’s right to understand how a grade was reached.


6. Backward Compatibility

Nothing existing changes. Scoring, attempt selection, and required/optional reporting are specified exactly as they behave. No score is recalculated, no grade moves, and no client sees a different value for anything it reads today.

Additions are optional. A client that ignores the new properties continues to work unchanged. New fields are additive to existing responses; no field changes meaning or type.

Historical work has no conditions, and must not be given defaults. Existing attempts carry no effort time and their placements carry no assistance tier. These are not recorded, not zero and not the lowest tier. Defaulting historical placements to scaffolded would cap the apparent attainment of every student in the platform’s history — a silent, large, and irreversible misrepresentation. Absent conditions are reported absent, and anything consuming them treats the corresponding evidence as unavailable rather than negative.

The two new conditions differ in whether they can be backfilled. Novelty is derivable retroactively, because the attempt history it rests on already exists — historical attempts can be resolved without new information. Effort time cannot: it was never observed and cannot be reconstructed. Fluency evidence therefore begins accumulating only from the point recording starts, and any measure depending on it must tolerate its absence for existing students indefinitely.

Existing tie behavior is being made consistent, not changed in intent. Two implementations of attempt selection currently disagree on fully tied attempts. §3.4 states the intended behavior so reconciliation has a specification to converge on rather than a choice to make.


7. References

Platform Documentation


Author

ĀYŌDÈ Development Team Codermerlin Academy Architecture