The agreement layer for a 23-campus system.
CSU's student-data problem is not a technology gap. It is the cost of getting 23 campuses to agree on what data means — one decision at a time. This document scopes a three-week engagement that makes each of those agreements cheaper — one decision at a time — starting with a proof anyone can check.
This page is our current understanding of the problem and the shape of a fix. It asks nothing of any campus. It exists to be corrected.
the problem isn't the data — it's the agreeing
Where it actually breaks.
Semantics
The same fact means different things on different campuses. Nine encodings of one term — and until governance conforms to one standard, the same question returns a different answer.
It runs deeper than codes: campuses hold their own versions of the same tables, with the same columns sometimes used for different things.
Decision
No mechanism exists for 23 campuses to agree. Every standardization is a political negotiation with no neutral evidence base and no named sign-off.
Decay
Finished projects don't stay finished. The crosswalk is a completed project with no maintenance owner. Drift resumed the day it shipped; the next field starts from zero.
so don't ask anyone to change — change where meaning lives
No campus is ever asked to change anything.
This is not a merge. No campus data is consolidated, and no university is asked to adopt another's standard — the cross-reference is the product.
FA26, F226, and F26 all remain valid at the source, forever. Campus Solutions instances remain the systems of record.
The overlay is a common language and a cross-reference — a system of meaning. The canonical definition exists only there — and each campus ratifies only its own mapping into it. Adoption costs a campus nothing and takes nothing away from anyone.
here is what that looks like in your environment
The system.
One chip per campus — each ratifies only its own mapping row.
Systems of record, never touched.
Matching and quality rules start from best practice and are derived from CSU's own policies once ingested — your documents become the standard the agents check against.
Reads schemas, the existing crosswalk, glossary exports.
Watches for new variants against ratified definitions; drift becomes a new recommendation, not a surprise.
Drafts each canonical definition with its evidence and the alternatives it rejected.
The system of meaning. Canonical definitions live here. Nothing above or below changes.
Each ratifies only their own mapping row. Nobody signs for the system.
Convenes the process; never signs for the campuses.
System-wide reporting · dashboards · (future) de-identified IR datasets.
This band is what the Decision Records section shows one entry of.
Fig. 1 — your estate flows as it always has. Agents read metadata and draft. Humans ratify. The overlay remembers.
and here is the loop that runs inside it
How an agent thinks its way through your catalogs.
The flow below is the actual method — shown here running on two CSU universities' published Accounting catalogs so every step has a real example attached. The point is the reasoning; the numbers are just what it looked like on one subject.
CSU Dominguez Hills and CSU Channel Islands · catalog year 2026-2027 · run 2026-08-17. Real extraction from public course catalogs. No system access, no credentials, no non-public data.
CSU Dominguez Hills
CSU Channel Islands
ACCT 310 · Intermediate Accounting · 3 units
The flow below runs once per course — 280 times for one subject, in seconds. Volume is the part that doesn't need a person.
Code → level and sequence position. Title → subject vocabulary. Description → topic vocabulary, with catalog filler stripped. Units → workload shape.
Generic words like 'course', 'introduction', 'topics' carry no meaning across catalogs, so they're removed before comparing.
No shortlisting, no sampling — every pair is scored, and every score keeps its evidence: how much the titles agree, how much the descriptions agree, whether the units line up, whether the course levels are near each other.
Titles are the strongest cross-catalog signal; descriptions confirm; units and level break near-ties. The weights are visible and adjustable — not a black box.
The agent drafts the pairing with its evidence and confidence attached.
Real example: ACCT 210 → ACC 230 · Financial Accounting on both sides · 0.739
Still reviewed — confidence sets the order of review, never skips it.
The agent presents the tie as a question, not an answer.
Real example: ACCT 310 tied at 0.68 against Intermediate Accounting I and Intermediate Accounting II — routed to a person, with both options and all evidence attached.
Tied at exactly 0.68 against both ACC 330 Intermediate Accounting I and ACC 331 Intermediate Accounting II. One university splits across two terms what the other teaches in one. This is a curriculum judgment, not a scoring problem.
No candidate clears the floor, so the course is recorded as unmatched — visibly, with the scores that failed.
Real example: ACCT 490 Special Topics found nothing — best score 0.214, below the >= 0.30 floor. Its true counterpart hides behind generic vocabulary.
A recorded miss can be caught. A silent one can't.
Proposal, tie, or absence: each outcome is written down with its evidence, its confidence, and a slot for the named human who resolves it. Nothing exits the flow undocumented.
Every human resolution is retained. When a person picks Intermediate I over II, the next near-tie in that subject starts from that precedent — the weighing gets smarter with every decision your people make.
Three places the machine could not finish the job.
Tied at exactly 0.68 against both ACC 330 Intermediate Accounting I and ACC 331 Intermediate Accounting II. One university splits across two terms what the other teaches in one. This is a curriculum judgment, not a scoring problem.
Ranked ACC 540 (graduate) above ACC 430 — same number, same title, undergraduate — by 0.02. A person resolves this instantly; the scoring did not.
Returned no match, though ACC 595 Selected Topics in Accounting exists. Both titles reduce to generic vocabulary once catalog boilerplate is stripped, so the signal disappeared. Recorded as a miss, not hidden.
This is the argument for the whole approach. The agents did the exhaustive part — 280 pairwise comparisons in seconds, every one with its evidence attached. What they could not do is decide. That is the part that stays with your people, and it is a much smaller part than it used to be.
The method does not change. Only the volume does — and volume is the part that does not need a person.
every one of these decisions leaves a record
How the agreement layer works.
Observe
Read metadata and the existing crosswalk — ground truth CSU already owns.
you see: every variant we found, listed
Recommend
AI drafts the canonical definition with its evidence and the alternatives it rejected.
you see: the draft definition, its evidence, its confidence score, and what we rejected
Ratify
The step that leaves the agents and enters the humans.
you see: your campus's row, and only yours, awaiting your name
Record
Who agreed, to what, why, and under what conditions — one lookup, forever.
you see: who agreed, to what, and why — one lookup, forever
Maintain
Drift detected against the ratified definition; changes re-enter at Recommend.
you see: drift flagged the day it appears, not in a dashboard later
AI arbitrates facts. Humans arbitrate values. The record is the product.
every pass through the loop leaves one of these behind
Where this goes.
First proof — small, bounded, PII-free
A test anyone can check, on data that raises no flags. Two candidate shapes:
a measured result — coverage and accuracy reported against work you already trust, every match carrying its evidence.
Term at scale
Canonical Term in the overlay across all 23 campuses, with distributed ratification by real campus representatives. Runs in CSU's AWS.
one field the whole system agrees on.
Governed expansion
Student/Person slice, FERPA-tagged, PII masking — the stepping stone to de-identified cross-campus IR datasets.
the same agreement machinery on governed, FERPA-tagged data, in your AWS
The 23-campus agreement layer
Every new definition enters as a decision record; drift is detected, not discovered in a dashboard.
a standing process: definitions enter as drafts, exit as agreements
The mechanics above are not a proposal — they run in production elsewhere:
A 350+ facility logistics operator
Unified data dictionary with full lineage, ontology-first. Facilities were never homogenized.
A global storage manufacturer
3,000+ tables from 10+ enterprise systems into a governed medallion architecture on AWS. Each table reviewed and approved by the client's own data team; seven quality dimensions and per-row provenance; run inside the client's own AWS account.
No campus changes anything. The system finally agrees.
Where we'd start: public or reference data only, nothing deployed, nothing changed — graded against work your team has already done.
