Programs and Workbench
Orchestrated (jobAsync) program challenges: the generator, runtimes,
comparators and scoring, workbench bundles and their authoring-time refusals,
Test in Workbench, and challenge-shell startup scripts. Verified against POC
build 213 (20260910T123348).
Design source: ARFC-1020, backend Epic #4644 with issues #4654, #4739, #4746, #4759, #4761, #4828 and #4829.
Reserved names and numeric bounds are provisional. Every reserved filename and every numeric limit on this page is marked PLACEHOLDER pending ARFC-1020 product confirmation in the platform source (
workbench_bundle.go), under issue #4739 — “do not treat any value below as final”. They are recorded here because the Studio mirrors them client-side and an implementer needs to know what it currently enforces, not because they are settled policy.
Contents
- What a program challenge is
- The generator
- Runtimes and the submission file
- Comparison and scoring
- Authored comparators
- Workbench bundles
- Bundle upload and digest gating
- Challenge-shell startup scripts
- Test in Workbench
What a program challenge is
A jobAsync variant of a challenge. Rather than comparing a typed answer
against a key, the platform runs the student’s program:
- The instructor authors a generator, bundled as a zip and uploaded at publish.
- In the job sandbox the generator runs and writes two files:
./input, which the student’s program receives on stdin, and./expected, the answer key. - The student’s program is linted, compiled and run against
./input. - A comparator scores
./outputagainst./expected, line by line.
Because the generator runs per attempt, every student gets different data from the same authored challenge.
The generator
A bundle of text files with a fixed entry point, run.sh, which runs with no
arguments and no stdin inside the sandbox and must write ./input and
./expected into the working directory. A new program challenge is seeded
with a working POSIX-sh template that generates twenty random integers and
their primality verdicts.
run.sh cannot be renamed or removed in the file list — it is the entry
point. Other files may sit beside it. Contents are text-only in v1:
ARFC-1020 §3.8’s reasoning is that admitting binary content to every channel
at once would widen three authoring surfaces for one demonstrated need.
The authoring surface validates continuously, showing problems as the instructor types and blocking publish on the same set:
| Refusal | Condition |
|---|---|
the generator run.sh is empty |
No content |
run.sh must write ./input and ./expected |
Neither literal path appears in the script |
The second check is deliberately textual — it cannot prove the script writes those files, only that it mentions them. It catches the common omission without pretending to be a sandbox.
Runtimes and the submission file
| Label | language/toolchain |
Extensions | Default file |
|---|---|---|---|
| C++ (g++, C++17) | cpp/gcc |
cpp, cc, cxx |
main.cpp |
| Python 3 | python/python3 |
py |
main.py |
The submission file name follows the chosen language automatically unless the instructor renamed it, so switching runtimes does not silently discard a deliberate name.
| Refusal | Condition |
|---|---|
submission file name must be a single file name (letters, digits, . _ -) |
Fails ^[A-Za-z0-9][A-Za-z0-9._-]*$ — no paths, no leading dot |
submission file should end in .cpp/.cc/.cxx for … |
Extension does not match the selected runtime |
Comparison and scoring
The standard comparator works line by line, in one of four modes:
| Mode | Compares |
|---|---|
| Boolean per line | true/false, 1/0, yes/no |
| String per line | Exact string |
| Integer per line | Exact integer |
| Numeric per line | To n significant figures |
Numeric mode requires 1–15 significant figures, refused outside that range.
Scoring is settled: the attempt passes only when every line matches, and
awardedScore = matched / total lines. Partial credit is therefore a real
signal about how much of the problem the student solved, not a
participation score.
Authored comparators
An optional bundle that, when present, replaces the standard comparison
entirely (ARFC-1020 §3.8). It keeps the generator’s conventions — a fixed
entry point, compare.sh, with other files allowed beside it, text-only in
v1.
An empty comparatorFiles list means “use the standard comparison”; that is
what makes the comparator optional rather than defaulted.
| Refusal | Condition |
|---|---|
comparator entry point must be "compare.sh" |
First entry is named anything else |
the comparator compare.sh is empty |
Entry point has no content |
Workbench bundles
Unlike generators and comparators, a workbench bundle can be attached to any challenge kind, not just programs (ARFC-1020 §3.3/§3.9/§3.10). It is the set of files that land in the student’s workbench directory — a starter file, scaffolding, data, reference material.
Workbench entries may be text or binary. Binary attachments are reported by name, type and size rather than edited, per §3.3: “the authoring tool reports their name, type, and size rather than pretending to edit them”. Unlike the generator, no entry is a fixed entry point — every entry is equally ordinary.
Reserved names
Refused at authoring time, matched on the basename, case-insensitively —
mirroring workbench_bundle.go’s checkWorkbenchBundleReservedName, which
remains the enforcement source of truth. Three exact names plus one pattern
(the platform source counts these as “four reserved names”):
| Reserved | Why |
|---|---|
question.md |
The platform’s own question-statement file |
.workbench-directory.json |
The platform’s directory record |
.workbench-config.json |
The student’s configuration overrides |
*.preserved, *.preserved.N |
The preserved-copy form |
Limits
Measured as the instructor assembles the bundle and reported live, never merely declared — again mirroring the server, which enforces them:
| Limit | Value |
|---|---|
| Stored (as-uploaded) archive | 64 MiB |
| Expanded size | 256 MiB |
| Entry count | 5,000 |
Hazards versus refusals
The Studio distinguishes the two, and the distinction is load-bearing: refusals block publish, hazards never do. The two hazards are things that are usually mistakes but legitimately intentional sometimes, so the tool warns and defers to the author:
| Hazard | Why it might be deliberate |
|---|---|
| The bundle places no file matching this program challenge’s declared submission file | The student may be meant to create it from scratch |
| No lesson terminal points at this challenge | Its files land in a directory the student has no terminal route to — though they can still merlin locate challenge |
Bundle upload and digest gating
All three bundle kinds are zipped in the browser using the store method
with a hand-rolled CRC-32 — the Studio has no build step and no bundler —
and uploaded through the synchronous execution-content endpoint (≤ 4 MiB;
content-addressed and idempotent by digest). The returned sha256 goes onto
the variant (for example generatorBundleDigest).
Re-uploads are gated locally: the Studio hashes the zip bytes and, if the hash matches what it uploaded last time, re-publishes with the remembered digest and performs no network write. An upload that returns anything other than a 64-hex-character digest is a hard error naming the challenge, never a silent pass.
Challenge-shell startup scripts
The authoring option :challenge-shell: <challengeRef>, written by
Insert ▸ Merlin Terminal ▸ Challenge,
is a Studio-level option the published component ignores. At publish the
Studio turns it into a real terminal configuration:
- Generate a startup script per referenced challenge that locates the
challenge directory via
merlin locate challenge --ref <ref>,cds into it, and reports either “Workbench ready in …” or an actionable message if the directory is not ready yet — most often because the lesson is not assigned to the student’s section. - Write it into the publish realm’s publisher zone, beside the published
lesson tree, at
lessons/published/<lesson>/shell/startup-<ref>.sh. - Rewrite the directive option into the component’s real options:
:startup-script: <path>@<version>plus:startup-script-owner: <publisherUserEID>.
Three implementation notes:
- The gateway sources the script under
set -euo pipefailbeforeexecing the login shell. That is why every command in the generated script is individually guarded and why thecdsticks. - Startup-script writes must not set a Content-Type.
writeAssetV2requires itsapplication/octet-streamdefault and 400s on anything else; the stored media type is sniffed from content, not from the request header. - A ref that matches no challenge fails the publish, with a message telling the author to re-insert from Insert ▸ Merlin Terminal ▸ Challenge — rather than publishing a terminal that would open nowhere.
Test in Workbench
Instructor self-test of a program challenge through the real
generate/run/compare and workbench pipeline — the unmodified publishLesson
path — against an isolated Instructor Preview sandbox that is never a real
section and never visible to students. It then mounts a live Merlin Terminal
on the freshly pinned startup script, through the same {merlin-terminal}
mechanism a real publish uses.
The preview environment
Membership is JIT self-enrollment, not a configured realm EID: every run calls the self-enroll endpoint asserting the Private realm, which grants the caller Instructor and Student membership if needed and echoes back the realm and role EIDs to use. It is idempotent, so repeat runs are cheap no-ops.
The scaffold is created-or-reused inside that realm — a subject, course, term
and section all coded instructor-preview. The preview lesson is then
released from draft before assignment, because from draft the only legal
transition is to released, a no-op PATCH is a 400, and a lesson may never
return to draft; a repeat test finds it already released and does nothing.
The scope limitation
Test in Workbench republishes the whole current lesson, not just the
challenge under test. publishLesson has no per-challenge entry point, and
fabricating a synthetic single-panel manifest by hand would risk diverging
from the published-lesson schema the real path already writes correctly.
The practical effect: every challenge in the lesson must pass publish validation — unique names, no empty questions, valid generators — before any one of them can be tested. A per-challenge-only preview would need a dedicated backend path; if the current restriction proves too coarse in practice, that is the follow-up.