ARFC-1012: Grading Policy and Rollup

Status

Draft Draft Date: 2026-08-08 Last Call Date: — Publication Date: — Version: 0.1

Abstract

A child who scores ninety-five percent on a single test may simply have had a good day. A child who has genuinely mastered a skill can perform it consistently, in unfamiliar situations, after time has passed, and can explain what they are doing. Mastery is not performance, and a platform that reports only performance cannot tell the two apart.

This ARFC specifies how recorded evidence becomes a reported result. It establishes two ways of doing that — a traditional score model that computes points and renders them on a scale, and a mastery model that judges whether a competency has actually been demonstrated — and it makes the second the platform’s recommendation while keeping the first fully supported. Both read the same evidence. Neither is derived from the other.

The mastery model rests on a single insight: the five levels a learner passes through are points on a gradient of decreasing help. Demonstrating something with substantial support, then in familiar situations, then unaided in expected contexts, then in novel ones, is one axis — and combined with whether the learner has met the work before, it produces a judgement that a percentage cannot express.

An institution configures which model applies, at whatever level of its structure the decision belongs to, using scales and matrices it may define itself.


Table of Contents

  1. Introduction
  2. Motivation
  3. Specification
  4. Version 1 Scope
  5. Security Considerations
  6. Backward Compatibility
  7. References

1. Introduction

ARFC-1010 specifies what the platform records about a student’s work: a score, how long they worked, how much the surroundings helped, and whether they had met the work before. ARFC-1011 specifies what that work is evidence for — which learning objectives it speaks to, and what kind of evidence it provides.

This ARFC specifies what that evidence is worth, and how it reaches a report card.

Two audiences need different answers from the same record. A registrar needs a figure for a transcript. A teacher deciding what to do on Monday needs to know whether a particular child has actually grasped two-digit addition or has been getting the right answers from the worked example still visible on the page. A single percentage serves the first badly and the second not at all.

So the platform computes both, from one evidence record, and an institution decides which one is authoritative for its reporting.


2. Motivation

Current State

The score model exists and is in production. A challenge’s contribution is its proportion correct multiplied by what the placement is worth; one attempt per challenge counts, chosen by the instructor; required and optional totals are reported separately. This is live in development and production.

Nothing else exists. There is no concept of a grading policy, no notion of a scale, no letter grades, no pass/fail, and no mastery judgement of any kind. Every institution on the platform receives points and percentages, whether or not that is how it reports.

“Objective met” has never been defined. Existing work on objective progress is explicitly blocked on it, recording that the product definition of what should be counted has not been settled. This ARFC is that definition.

Missing Capabilities

An institution cannot express how it grades. Schools that report letter grades, schools that report pass/fail for particular cohorts, and schools that report standards-based proficiency all receive the same percentages and must convert them elsewhere, outside the platform and outside its audit trail.

Performance cannot be distinguished from mastery. Because the conditions of an attempt were never recorded, no distinction was possible. ARFC-1010 makes it possible; nothing yet acts on it.

A single high score ends the enquiry. There is no way to express that a competency should be demonstrated more than once, in more than one kind of situation, or after time has passed — which is what teachers mean when they say a child has mastered something.

Early failure is permanent. With points-based grading, a low mark from a first attempt at something a student later masters completely remains in the average forever. The record reflects the timing of learning rather than its eventual level.

No cohort can be graded differently. A group of students who should receive pass/fail rather than percentages cannot be accommodated, so the accommodation happens off-platform.

Design Goals

  1. Recommend mastery, support score. Mastery is the platform’s default. Score-based grading remains fully supported, first-class, and never second-rate.
  2. One evidence record, two readings. Neither model stores its own evidence, and neither is computed from the other’s output.
  3. Make the two impossible to conflate. The design must prevent a percentage from becoming a level or a level from becoming a percentage by convenience.
  4. Let students recover. Early attempts at something later mastered must not permanently depress the result.
  5. Let institutions express themselves. Scales and judgement matrices are configurable, not hard-coded, and apply at whatever level of an institution’s structure the decision belongs to.
  6. Explain every judgement. Any reported level can be traced to the specific evidence that produced it.

3. Specification

3.1 Core Principles

Mastery is not performance. A score measures a performance on one occasion. Mastery is a claim about a capability, and a capability is shown by repeated demonstration under varying conditions. The two require different treatment because they are different claims.

Sufficiency, not averaging. A mastery judgement asks whether sufficient evidence exists that a competency has been demonstrated. It does not average scores, and it cannot, because levels are ordered categories rather than quantities.

Every judgement is explainable. A teacher asked in a parent conference why the platform says a child is proficient must be able to answer. The evidence behind a level is recorded, not recomputed from scratch and hoped to match.

Configuration is data, not code. The recommended levels, scales, and matrix ship as ordinary configuration that an institution may replace. There is no privileged built-in behaviour that customisation works around.

Proficient is the target. Advanced exists to recognise achievement beyond what was required. It is reported and never required, which is what makes it safe to have on the scale at all.

3.2 The Integrity Rule

Two operations are prohibited. They are stated here, before anything else in the specification, because both are convenient, both will be proposed, and either one reduces mastery reporting to decoration on a gradebook.

A mastery level must never be derived from a percentage. Ninety-five percent is not proficiency. A student who scores ninety-five percent once, on a familiar problem, beside a worked example, with unlimited retries, has demonstrated nothing about whether they can do it unaided or in an unfamiliar situation. Converting a score into a level manufactures a claim the evidence does not support.

A percentage must never be derived by averaging levels. Levels are ordered categories, not quantities. The distance from Emerging to Developing is not a measurable interval, still less the same interval as Developing to Proficient, so a mean over them is undefined. Where a single figure is required, it is produced by counting objectives that meet a threshold (§3.15), never by averaging.

Both prohibitions are properties of the platform, not settings. No configuration enables either.

3.3 Primary Objects

Grading model — either score or mastery. Determines how evidence becomes a result.

Scale — an ordered set of named bands over a domain, which turns a computed result into something reportable. Percentages, letter grades, pass/fail, and the mastery levels are all scales.

Demand matrix — for the mastery model, the mapping from the conditions under which a student succeeded to the level that success supports.

Grading policy — a model, a scale, and (for mastery) a matrix, applied to some part of an institution.

Mastery record — for one student and one objective, the level reached, the evidence that produced it, and the policy in force when it was judged.

3.4 Two Models, One Evidence Record

flowchart TB
    EV[Recorded evidence<br/>score · effort time · assistance · novelty<br/>aligned to objectives]
    EV --> SM[Score model<br/>proportion × worth, summed]
    EV --> MM[Mastery model<br/>conditions → level per objective]
    SM --> SC[Rendered on a scale<br/>87% · B− · Pass]
    MM --> ML[Rendered as levels<br/>Inevident → Advanced]
    SC --> RPT[Reported result]
    ML --> RPT
    POL[Grading policy<br/>selects which is authoritative] --> RPT

Both models remain computable wherever their inputs exist, and the policy selects which is authoritative for reporting. The other stays available as a secondary view, because most institutions using mastery internally still maintain a gradebook.

The two are not equally available, and this asymmetry is why the score model must remain first-class:

  Score model Mastery model
Availability always — every placement already carries what it needs requires investment
Additional authoring none objectives, alignment, assistance tiers, item variety, policy

A course that has not established objectives or authored assistance tiers cannot produce mastery levels, and must report that they are unavailable rather than emit levels computed from absent data. Score is the floor that always works; mastery is earned.

3.5 Mastery Levels

Level Meaning
Inevident No observable evidence of the competency.
Emerging Beginning to demonstrate the competency, with substantial support.
Developing Demonstrates the competency inconsistently, or only in familiar situations.
Proficient Demonstrates the competency independently and consistently in expected contexts. This is the target level.
Advanced Transfers the competency to novel situations, explains their reasoning, or demonstrates exceptional sophistication.

Read the wording of the middle four and a single axis appears: with substantial support, then in familiar situations, then in expected contexts, then in novel situations. The levels are points on a gradient of decreasing help. That is what makes them computable from recorded conditions rather than from a teacher’s impression, and it is the foundation of §3.6.

Inevident is not a low score. It means nothing has been observed. A student who has attempted nothing and a student who has attempted everything and failed are in entirely different situations, and both are misrepresented by the same label.

3.6 The Demand Matrix

Two independent conditions from ARFC-1010 describe how demanding a demonstration was.

Contextual assistance — how much the surroundings helped: scaffolded → practiced → standard → unaided.

Novelty — whether the student had met the work before: repeat → reformed → novel.

A student’s level for an objective is the most demanding combination at which they have succeeded accurately and consistently. The recommended mapping:

  repeat reformed novel
scaffolded Emerging Emerging Emerging
practiced Emerging Developing Developing
standard Developing Proficient Proficient
unaided Developing Proficient Advanced

The two bold cells in the lower left are the reason both axes are needed. A student answering a question they have already met, even under unaided exam conditions, has shown recall — not competence. A single combined measure of “difficulty” would credit that as Proficient or Advanced; separating the axes caps it at Developing. Any institution reusing items between practice and assessment would otherwise see levels inflate systematically.

The matrix is configuration. An institution may replace it. Two constraints hold for any matrix:

Inevident is not a cell. It is what holds when no combination has been satisfied.

3.7 Accuracy and Consistency Gates

Reaching a cell requires more than encountering work under those conditions.

Accuracy — the student’s result for the counting attempt must meet the objective’s accuracy threshold, expressed as a proportion of what the work was worth. The threshold is part of policy, not fixed by the platform.

Consistency — a single success does not establish a level. An objective’s policy states how many qualifying successes are needed, and the platform’s recommended default is more than one. This is what the word consistently in the definition of Proficient means, made operational.

Because consistency requires several successes, a level is a statement about a pattern of evidence rather than about any one occasion — which is precisely the distinction between mastery and performance.

3.8 Additional Requirements per Objective

Beyond the two demand axes, three further conditions may be required for an objective, because they are not equally relevant to every competency.

Requirement Satisfied when Comes from
Retention the objective is evidenced again after a defined interval has passed the timing of the evidence
Explanation evidence exists from work that asserts explanation for this objective the alignment’s assertions and the challenge’s kind
Fluency the student worked within the expected duration for the work’s kind recorded effort time

Whether each is required is set per objective. This matters because competencies genuinely differ: good digital citizenship should be demonstrated across a year and in several situations, while Boolean conjunction may be settled thoroughly in one lesson and needs no spaced re-check. A single global rule would be wrong for one of them.

Where a requirement is set but its evidence is unavailable, the objective does not reach the level that required it, and the report says which requirement is outstanding. It never silently passes, and it never fails as though the student had performed badly.

3.9 Rising and Falling

A level rises freely. Any qualifying evidence that supports a higher level raises it, whenever it arrives. A student who struggled in September and demonstrates the competency thoroughly in March is proficient; the September attempts do not hold them down. This is the recovery property that points-based averaging cannot provide, and it is why mastery reporting reflects the eventual level of understanding rather than the timing of learning.

A level falls only on newer evidence against a time-sensitive requirement. If retention is required and a spaced re-check fails, that failure is meaningful and the level reflects it. Nothing else lowers a level: no later attempt at easier work, and no accumulation of poor attempts alongside good ones.

The asymmetry follows from what the levels claim. Can they do this? is answered by their best demonstration. Can they still do this? is answered by their most recent one. The first can only improve with more evidence; the second genuinely can regress.

3.10 Scales

A scale is an ordered set of named bands over a domain. Every reportable form of a grade is a scale:

Scale Domain Bands
Numeric percentage none — the value is reported directly
Alphabetic percentage as many as the institution defines, e.g. A+ through F
Pass/fail percentage or level two
Mastery levels level the five of §3.5

Treating all four as one concept is what makes custom scales inexpensive. An institution defining its own letter-grade cutoffs, or its own two-band pass threshold, is defining bands — not requesting a feature.

Each band carries a label and whether it counts as passing. Two constraints hold:

Scales are shared configuration: the platform ships the recommended set, and an institution adds its own without replacing them.

3.11 Grading Policy and Where It Applies

A grading policy is a model, a scale, and — for mastery — a matrix. It applies to a part of the institution.

An institution nests: a district contains schools, a school contains classes, and a class may contain a smaller group. A policy attaches at any of those levels, and the practical case that requires the smallest is real: a specific group of students who should receive pass/fail rather than percentages, most often for accommodation reasons.

Two properties are normative.

The group a policy attaches to must be deliberately constituted. A grading policy may not attach to a cohort the platform assembles automatically. The platform can group students by inferred characteristics — a readiness band, a shared language, an accommodation — and those groupings are recomputed as students change. A grading policy attached to one would silently move a student between grading schemes when a recomputation reclassified them. Grading designation is a formal, often legally significant decision, and it attaches only to a group someone deliberately created.

An individual may be named directly. Where a single student requires a different scheme, they are named, in the same way individual accommodations already work elsewhere in the platform. This avoids constituting a group of one.

3.12 Policy Resolution

flowchart TB
    P[Platform default: mastery] --> D[District policy]
    D --> S[School policy]
    S --> C[Class policy]
    C --> G[Group policy]
    G --> I[Named individual]
    I --> R([Policy in force])
    style P stroke-dasharray: 4 4

The policy in force for a student is the most specific one that applies: a named individual overrides their group, which overrides their class, which overrides the school, then the district, then the platform default.

Because an institution’s levels nest strictly — each part has exactly one parent — there is always exactly one most specific policy. No ambiguity can arise, and this is a further reason a policy may not attach to an automatically assembled cohort, which could span several classes and produce two competing answers.

Granularity and authority are separate questions. That a policy may be set on a single class does not mean every teacher should set one; each teacher choosing their own scale would make comparison across a school meaningless. The control is who may set a policy at each level (§3.18), not whether the level exists.

3.13 The Mastery Record

For each student and objective, the platform records the level reached together with everything needed to explain it:

Recording why work did not count is as important as recording what did. “Not enough evidence” is unanswerable in a parent conference; “three demonstrations qualified, two more were under scaffolded conditions and one is awaiting a spaced re-check in three weeks” is a conversation a teacher can have.

Because the policy in force is recorded alongside the judgement, a level remains interpretable after a policy changes. A grade reported in October under one scheme does not become unexplainable when the district adopts another in March.

A judgement is superseded, never overwritten. Re-judging records a new assessment and marks the previous one superseded, so the history of a student’s reported attainment is auditable.

3.14 Freshness and Recomputation

A mastery judgement depends on three things that change independently: the evidence beneath it, which objectives that evidence is aligned to, and the policy in force. When any of them changes, affected judgements are re-made.

Judgements are stored and refreshed, not recalculated on demand. A report therefore reflects a known point in time, and every report states when its judgements were made. This is the same principle the platform already applies to competitive scoring, where an authoritative record of events is the source of truth and the aggregates over it are maintained and recomputed rather than derived afresh on every read.

The consequence for consumers: a judgement may briefly lag a regrade. A report must show its own freshness rather than implying it is live, and staleness must be visible rather than inferred.

3.15 Reporting a Single Figure

Where a transcript requires one number from a mastery course, it is produced by counting the required objectives that reach the target level, as a proportion of all required objectives.

This is a count, not an average, which is what keeps it compatible with §3.2. It also states something a reader can interpret without knowing the scheme: how much of what this course required has this student demonstrated?

Optional objectives are reported separately and never contribute to the figure, exactly as optional work is kept separate in the score model.

Objectives with no evidence must not be counted as failures. An objective a student has not yet had the opportunity to demonstrate is not a shortfall on their part. A figure that silently treats Inevident as zero reintroduces precisely the distortion mastery grading exists to escape, and it misrepresents a course that has not finished teaching as a student who has not learned. Where required objectives remain unevidenced, the figure is reported as provisional, over the objectives actually available.

3.16 What the Plan Must Provide

Mastery requirements impose obligations on a course’s plan, which become checkable before a term begins rather than discoverable at its end. A plan is checked for whether it can actually deliver the target level:

Each is reported as a planning defect with the objective named and the missing condition stated — “this objective is only ever assessed with work students have already practised” — rather than as a generic warning. ARFC-1013 specifies where these checks surface in the planning experience.

3.17 API Surface

Method Path operationId Description
GET /v2/realms/{realm-eid}/grading-policy getRealmGradingPolicyV2 The policy set at this level, and the effective policy after resolution
PUT /v2/realms/{realm-eid}/grading-policy setRealmGradingPolicyV2 Set or replace the policy at this level
DELETE /v2/realms/{realm-eid}/grading-policy clearRealmGradingPolicyV2 Remove this level’s policy so the parent’s applies
PUT /v2/realms/{realm-eid}/grading-policy/students/{student-eid} setStudentGradingPolicyV2 Name an individual receiving a different policy
GET /v2/grading-scales listGradingScalesV2 Platform-provided and institution-defined scales
POST /v2/realms/{realm-eid}/grading-scales createGradingScaleV2 Define a scale, with its bands
GET /v2/mastery-matrices listMasteryMatricesV2 Platform-provided and institution-defined matrices
POST /v2/realms/{realm-eid}/mastery-matrices createMasteryMatrixV2 Define a matrix; rejected if incomplete or inverting
PATCH /v2/courses/{course-eid}/objectives/{objective-eid}/mastery-requirements updateObjectiveMasteryRequirementsV2 Set accuracy threshold, consistency count, and required additional conditions
GET /v2/sections/{section-eid}/students/{student-eid}/mastery getStudentMasteryV2 Levels per objective, with freshness
GET /v2/sections/{section-eid}/students/{student-eid}/mastery/{objective-eid} explainStudentMasteryV2 The full record: qualifying evidence, excluded evidence with reasons, outstanding requirements, policy in force
GET /v2/sections/{section-eid}/mastery getSectionMasteryV2 Section-wide levels per objective
GET /v2/courses/{course-eid}/plan-readiness getCoursePlanReadinessV2 Planning defects that would prevent the target level being reached

Setting a policy returns the resolved effective policy, not merely an acknowledgement, so a caller can see what a change actually produced for the students beneath it.

3.18 Authorization Model

Setting a grading policy is an institutional act. By default it requires authority over the part of the institution it applies to: a district administrator for a district, a school administrator for a school. An institution may delegate class-level and group-level policy to a curriculum lead. Instructors do not set grading policy by default, precisely because comparability across a school depends on their not doing so individually — though an institution may grant it.

Defining scales and matrices is separate from applying them, and more restricted. Composing a scale changes what every grade in its scope means, so an institution may allow a curriculum lead to select among defined scales while reserving defining them to an administrator.

Setting per-objective mastery requirements requires curriculum authority over the course, matching who may establish its objectives.

Reading a mastery record follows existing scope. A student may read their own and no one else’s. An instructor may read records for students in sections they are assigned to. The explanation of a judgement is available to whoever may read the judgement — a level without access to its basis would be unanswerable.

A change of policy is attributable. Every policy, scale, and matrix change records who made it and when, because each one can move reported attainment for many students at once.

3.19 Absent, Insufficient, and Not Demonstrated

Five states must remain distinct. Collapsing any pair into another is a defect, and the last three are the ones most often collapsed by mistake.

State Meaning Must not be shown as
Mastery unavailable the course has not established what mastery requires a level, or a zero
Inevident no evidence has been observed a failure
Insufficient evidence some evidence exists but not enough to judge Not demonstrated
Not demonstrated sufficient evidence exists and the level was not reached Inevident
Requirement outstanding the level is otherwise reached but a required condition is unsatisfied Not demonstrated

The difference between Insufficient evidence and Not demonstrated is the difference between “we do not yet know” and “we know, and not yet” — the first calls for more opportunity, the second for different teaching. A report that renders them alike removes the information a teacher needs most.


4. Version 1 Scope

Included

Excluded


5. Security Considerations

A grading policy change moves many students’ reported attainment at once. Replacing a scale or a matrix can alter every grade in its scope without any student’s work changing. Such changes are attributable, and because each judgement records the policy that produced it, a past report remains explicable afterwards. Authority over policy is therefore a materially different question from authority over an individual grade.

A matrix is the most dangerous configuration surface. A matrix that awards Proficient for scaffolded repeats would inflate every level in its scope while appearing to be ordinary configuration. The non-inversion constraint prevents the most obvious form of this; it does not prevent a uniformly lenient matrix, so matrix definition is restricted and auditable rather than merely validated.

The explanation of a judgement is a disclosure surface. It necessarily reveals which work a student attempted and how they performed. It is scoped exactly as the judgement is, and must never expose the evidence of other students — including through comparative framing such as how a student ranks within a cohort.

Levels are sensitive judgements about children. A mastery level is a more consequential statement than a score, since it purports to describe capability rather than performance on a day. Levels are not exposed to peers, are not aggregated into student-visible comparisons, and carry the same retention and disclosure obligations as grades — including the right of a student or parent to understand how a level was reached, which §3.13 exists to make possible.

Freshness is a correctness property, not a convenience. A stale judgement presented as current could support a consequential decision — placement, intervention, promotion — on superseded evidence. Reports state when judgements were made.


6. Backward Compatibility

No existing grade changes. The score model is specified exactly as it behaves. Every institution not adopting a policy continues to receive what it receives today, and the platform default applies only where mastery is actually configured.

The default is a recommendation, not an imposition. Although mastery is the platform’s recommended model, a course that has not established objectives, alignments, and assistance tiers reports mastery as unavailable and falls back to score. No institution is required to act, and none finds its reports altered by this ARFC alone.

Mastery cannot be computed retroactively in full. Historical work carries no assistance tier and no effort time (ARFC-1010 §6). Novelty is derivable from existing history, but assistance is not, and effort time was never observed. Adopting mastery therefore begins accumulating levels from the point the conditions are recorded, and a report covering an earlier period must state that its evidence predates condition recording rather than presenting incomplete levels as complete.

Absent conditions must not be given defaults. Treating a missing assistance tier as scaffolded would cap every historical demonstration at Emerging across the platform’s entire history. Absent means unavailable, and an objective whose evidence lacks the conditions to judge it reports insufficient evidence.

Both models coexist permanently. This is not a migration with an end state. Score-based grading is not deprecated, has no sunset, and remains the only model that works without additional authoring.


7. References

Platform Documentation


Author

ĀYŌDÈ Development Team Codermerlin Academy Architecture