ARFC-1019: Exact String Comparison and Value Normalization
Status
| Last Call | Date: 2026-09-06 | Version: 0.1 |
Last Call window: 48 hours (compressed from the standard 14–21 days), closing 2026-09-08. Compression reason: this ARFC gates time-boxed platform work (case-sensitivity remediation for #4707/#4713) running under active time pressure; the design itself is a narrow generalization of the already-drafted ARFC-1018 rather than novel scope. Precedent for a compressed lifecycle: ARFC-1008 ran Draft 2026-07-22 → Last Call 2026-07-24 → Published 2026-07-24; ARFC-1007 published the same day its Last Call opened.
Abstract
This ARFC states how the platform decides whether two pieces of text name the same thing: exactly. Two values match only when they are the same sequence of characters — case, accents and character width are all significant.
Where insensitivity to case is genuinely part of a value’s meaning, as with an email address or a username, the platform does not compare loosely. It records the value in one canonical form, so that an exact comparison produces the insensitive answer a caller expects. Searching stays deliberately insensitive and says so.
ARFC-1018 established this rule for pathnames. This ARFC generalizes it to every value the platform stores, and states the three classes of value the rule resolves into.
Table of Contents
- Introduction
- Motivation
- Specification
- Security Considerations
- Backward Compatibility
- References
- Author
Introduction
Almost every request names something by text — a path, an email address, a username, a course code, a search term. Each one implies an answer to a question the platform has never published: when do two spellings mean the same thing?
In the absence of a stated answer, different parts of the platform arrived at different ones, and the answer that took hold most widely was inherited rather than chosen. This ARFC states one rule, names the small number of exceptions, and requires that each exception be declared rather than assumed.
Motivation
Current State
Text comparison across the platform is insensitive to case and to accents, inconsistently. Verified against live behavior:
| Compared | Today |
|---|---|
A and a |
same |
a and á |
same |
a and â |
same |
ae and æ |
same |
o and ø |
same |
A and A (full width) |
same |
s and ß |
different |
i and ı |
different |
No author could state that rule from memory, and none of it is documented. The platform is also inconsistent with itself: two files whose names differ only in case are two files, while two directories whose names differ only in case are one directory.
Why This Is Not a Decision Anyone Made
The insensitivity is a side effect of a compatibility alignment performed to resolve a technical error, not a choice about what identity means. Where insensitivity genuinely was intended — matching an email domain, for instance — it was written explicitly and commented as such. Everywhere else it was simply inherited.
It has nevertheless become load-bearing in places nobody examined. An authorization decision matches a signed-in user’s email address against a recorded membership request; notification suppression matches a recipient address against a suppression list; automatic enrolment matches an email domain. All three work today only because the comparison quietly ignores case, and the values they compare are not recorded consistently.
That is the real hazard: an implicit rule that nobody chose, that nobody can state, and that several correctness- and authorization-sensitive behaviors now depend on.
Use Cases
- A caller can predict the answer. Given two spellings, an integrator can say whether the platform will treat them as the same thing, without consulting anyone.
- Work moves between systems that compare exactly. Student work moves between a case-sensitive shell and the asset store; published content is addressed by URL, where paths compare exactly. A platform that folds while its neighbours do not will eventually disagree with itself.
- Insensitivity is a choice, stated per value. An email address ignoring case is a decision about what an email address is, and should read as one — not as a property of where it happens to be stored.
- One rule a person can hold in their head. Treating
A,a,áandâas four distinct characters is such a rule. Conflating some pairs and not others is not.
Design Goals
- Exact comparison by default, everywhere.
- Insensitivity only where it belongs to the meaning of the value, achieved by recording one canonical form rather than by comparing loosely.
- Every value’s rule is declared, and discoverable from the platform itself.
- No silent substitution: a request is never satisfied by something the caller did not name.
- A violation of the rule is visible rather than absorbed.
Specification
Three Classes of Value
Every stored text value belongs to exactly one class, and the class is declared.
| Class | Examples | Rule |
|---|---|---|
| Address | pathnames, resource keys | Compared exactly; the caller’s spelling is preserved and returned |
| Normalized identity | email addresses, usernames | Recorded in one canonical form (lower case); a caller may present any case and receives a consistent answer; the platform returns the canonical form |
| Search term | the query parameter of a listing endpoint | Matching is deliberately insensitive, and the endpoint’s documentation says so |
Anything not declared as a normalized identity or used as a search term is an address: compared exactly.
Core Rules
- Exactness is the default. Two values match only when they are the same sequence of characters. Case, accents and character width are significant.
- Insensitivity is achieved by normalization, never by loose comparison. When a value’s meaning ignores case, the platform converts it to its canonical form at the moment it is recorded. Comparison remains exact.
- The canonical form is what the platform returns. A caller who submits
Instructor@School.eduand later reads the value back receivesinstructor@school.edu. - Every value’s class is declared. The platform can state, for any value it stores, which rule governs it. A value with no declaration is an address.
- Authorization comparisons are never loosely compared. A decision about what a caller may do compares addresses exactly and identities in canonical form. It is never satisfied by a value differing only in case, accent or width from the one named.
- Search is the stated exception. Listing endpoints that match a query fragment do so insensitively and document that behavior; this never extends to identifying or authorizing.
What a Caller Observes
| Action | Result |
|---|---|
| Read a path with different case than it was written | Not found |
| Create two paths differing only in case | Two distinct objects |
| Sign in or be matched by an email address in any case | Works — the recorded value is canonical |
| Read back an email address submitted with capitals | Returned in canonical (lower-case) form |
| Search a name with different case | Matches, as documented for that endpoint |
Declaring a Value’s Class
A value’s class is a property of the platform’s own description of itself, not of convention or of a reviewer’s memory. Two obligations follow:
- A normalized identity carries its declaration wherever the value is described, so that a reader encountering it knows the rule without inspecting behavior.
- The declaration is enforced, not merely recorded. The platform rejects a value that does not satisfy the class it declares, at the moment of writing. A normalized identity cannot come to hold a non-canonical value, whatever path put it there.
Declaration and enforcement together are what keep this ARFC from decaying into the situation it replaces: a rule that is true of the code as first written and progressively less true afterwards.
Version 1 Scope
Included
- Exact comparison as the default for every stored text value.
- The three classes, with each value’s class declared and enforced.
- Canonicalization of normalized identities at the moment of recording, and return of the canonical form.
- Exact comparison in every authorization decision.
Excluded
- Unicode normalization — whether two different encodings of the same visible character are the same value. Deliberately deferred, as in ARFC-1018, and a candidate for its own proposal.
- Renaming, merging or rewriting existing content, except the canonicalization of values in classes declared as normalized identities.
- Any change to how values are displayed or ordered.
- Any similarity detection: the platform does not judge whether two different values were meant to be the same, and does not warn or refuse when they look alike. A different value is a different value.
- Case-insensitive behavior in any class other than search.
Security Considerations
Silent substitution is removed. Under a loose comparison, a request naming one thing can be satisfied by another whose name differs only in case — a confusion that matters most where the name selects something privileged. Exact comparison eliminates the class.
Authorization is the sharpest case. Several authorization decisions compare text: a caller’s identity against a recorded grant, a requested path against the content a grant covers. Rule 5 exists because a loose comparison in that position is not merely inconsistent — it can authorize a caller against something they did not name. Any comparison used to decide access must be exact or canonical, and this ARFC treats a loosely-compared authorization predicate as a defect.
Normalization protects account identity. Exactness alone would allow two accounts whose usernames differ only in case, which is an impersonation surface. Canonical recording of identities closes it: the two cannot both exist, because they are the same recorded value.
Look-alike values remain distinct. Because more characters are now distinct, two values may appear nearly identical when displayed. Authorization depends on grants over objects, never on name similarity, so this is not an access-control risk — but an interface presenting values to a person should not treat visual distinctness as a safety property.
The change only narrows. Every request that succeeds under these rules would also have succeeded before; the rules remove matches, never add them.
Backward Compatibility
- Addresses become strict. A caller that previously reached an object by spelling its path differently now receives not found. This is the effect ARFC-1018 describes, generalized.
- Identities keep working. A caller presenting an email address or username in any case continues to be matched, because the recorded value is canonical. This is the case where behavior is deliberately preserved.
- Returned identity values may change case. A client that stored a copy of what it submitted and compares it byte-for-byte against what the platform returns should compare canonically instead.
- Search is unchanged.
- No content is renamed or merged, and no two existing values become ambiguous under these rules.
- Integrators addressing objects by URN are unaffected; URN resolution does not depend on these rules.
References
- ARFC-1018 — Case-Sensitive Asset Pathnames (the pathname instance of this rule)
- ARFC-1001 — URN-Based Asset Identifiers (case rules for URN identifiers)
- RFC 3986 — Uniform Resource Identifier: Generic Syntax (path comparison)
- RFC 5321 — Simple Mail Transfer Protocol (address case semantics)
codermerlin.academy-backend#4713— implementation of this ARFCcodermerlin.academy-backend#4707— the incident that prompted ARFC-1018
Author
ĀYŌDÈ Development Team