White Paper · Vertical Language Models

Beyond
the Big Model

Why enterprises need Metis: a private vertical model for Enterprise Intelligence and the judgment capital a company creates every day.

The scarce asset in enterprise AI is no longer access to a smart model. It is ownership of what the firm teaches through use.

Metis is a Vertical Language Model for enterprises: smaller, narrower, private to the firm, and closer to the business. It is shaped around the language of a market, the shape of a workflow, the policies around it, and the industry signals that make a company different.

The short version

Most enterprise AI roadmaps start by asking which big model to use. That is a good first question. It is not the last one.

The deeper question is where enterprise judgment accumulates. Every query, correction, accepted draft, rejected answer, evaluation, decision, and workflow trace teaches the system something about how the firm thinks. That is not exhaust in the ordinary sense. It is judgment capital.

This is the Judgment Capital Gap: the distance between where enterprise judgment is created and where it compounds. If a firm's best corrections, examples, evaluations, and decisions improve someone else's intelligence estate, the company is exporting the very advantage AI was supposed to amplify.

Metis is the enterprise answer: a Vertical Language Model where the company's language, standards, examples, and judgment compound inside its own boundary, close to the people and agents doing the work.

The future enterprise stack is not one giant model. It is a governed family of vertical models.

00 / The short version

The Judgment Capital Gap

Metis starts from a simple belief: the intelligence a company creates while doing its work should stay with the company.

Every enterprise has a private operating texture: how it ranks tradeoffs, what it considers a good answer, which market signal matters, which policy phrase is acceptable, and when a workflow should escalate. That texture is not generic data. It is judgment capital.

The Judgment Capital Gap opens when that capital is created inside the firm but captured outside the firm. The company gets useful answers today, while the pattern of its judgment can strengthen an intelligence system it does not own.

Metis closes the gap by giving the enterprise a private Vertical Language Model where its language, examples, standards, and repeated decisions compound inside its own boundary.

There is one line drawn inside the firm’s own material, and it is the line that keeps compounding from turning into drift: the model does not train on its own opinions. The test is not who typed the sentence but what the sentence asserts. A fact a research agent went out and retrieved, anchored to a source, is material the model can learn from. The model’s own judgment, ranking or rationale about that fact is not - a model that keeps re-reading its own conclusions drifts away from the world rather than toward it. Telling those two apart is itself a judgment, so it is made by a model rather than by a keyword list, and it fails closed: a row nobody can confidently classify is discarded, not waved through as probably fine.

01 / The problem

AI exhaust is core IP

The most valuable enterprise knowledge is often too lived-in to look like an asset on a balance sheet.

It appears as the partner's redline, the operator's exception, the investment committee's phrasing, the analyst's rejection, the risk team's escalation rule, and the compliance team's quiet standard for what counts as acceptable.

In the AI era, those traces become machine-readable. A firm therefore needs more than file security. It needs a model strategy that keeps the compounding value of its own judgment under its own control.

In Foundation that control is structural rather than contractual. A firm’s model is built from that firm’s own records and from material that belongs to no customer at all; another tenant’s rows are excluded by the query that assembles the corpus, not by a policy somebody has to remember to apply. One firm gets one model. A second firm gets a second model, under its own name, with its own corpus. That is what keeps confidential material from ever becoming cross-customer exposure - the model that consumed a firm’s data only ever serves that firm.

What is excluded is worth being precise about, because it is close to the opposite of what people expect. Sensitive and confidential material is not filtered out: for a model the enterprise owns, and which serves nobody else, that material is the entire point. What is filtered out is anything without governed provenance. Material may enter training only from the platform’s own recorded surfaces - the ones carrying an ingest receipt, a rights record and a content hash for every row. A clean-looking file that arrived from nowhere in particular is inadmissible precisely because nothing can say where it came from.

An approval to use a body of material also has to name what it covers: the owner, the store, the scope, the tenant. A blank field is refused outright rather than quietly filled in with a placeholder, because an approval that names nothing would approve everything.

A company should not rent intelligence by donating the judgment that makes it valuable.

02 / Metis

A Vertical Language Model for the enterprise

Metis is a small language model built vertically and kept inside the enterprise trust boundary. Every Metis call is pinned to a local-only address - any other endpoint is refused rather than quietly used - and before an answer is handed back, the runtime checks that the model which replied is the exact approved artifact, by digest, fully resident on the local GPU. If it is not, the call fails instead of returning, and no other model is allowed to stand in for it. In this paper, a Vertical Language Model means a model shaped around a specific enterprise domain rather than a general chatbot trying to stretch across every topic.

That vertical can be an industry, a function, a product line, a regulated workflow, or a company-specific operating pattern. What matters is that the model is designed and evaluated around the work it is meant to do.

Metis has five jobs:

  • Speak the company's language. Use the names, categories, exceptions, and shorthand that real teams use.
  • Understand the workflow. Know the steps, handoffs, constraints, and judgment calls inside a vertical process.
  • Stay inside the domain. Be excellent at the work it is responsible for, know when to hand off to a broader model or a human, and record what it did not know - so that once a gap is confirmed, the enterprise can send its own research agents to go and acquire the missing material.
  • Use trusted knowledge. Ground its behavior in the material, examples, and operating rules the enterprise actually relies on.
  • Serve enterprise agents. Give assistants, workflows, and teams a focused model they can call repeatedly without asking a generalist to rebuild the context every time. In Foundation this is not a specialist you call now and then. It is the model the platform reaches for first, across chat, operations, and synthesis, with the larger outside model kept for escalation.

This is why small matters. A smaller model can be cheaper to run, easier to evaluate, easier to specialize, and easier to govern. It becomes useful not by knowing everything, but by knowing the right thing deeply.

None of that is worth much unless the model actually improves, so it is worth saying plainly how a new version of Metis comes to exist. A round of work runs on a fixed cadence and moves through the same ordered steps every time: select the material, assemble it into a training set, train, evaluate, and only then adopt. Each step writes a receipt as it finishes, and the next pass reads those receipts before doing anything - so an interrupted round resumes where it stopped instead of starting over, and a round that has already trained is never quietly trained a second time. A short trial run has to succeed before the long one is allowed to start, and if a training run is already using the machine the next cycle stands down rather than starting a second one alongside it.

Adoption is where the discipline shows. A newly trained candidate is measured on material held back from its own training, and it has to clear three gates - two of them measured against whatever is currently in service, one an absolute standard it must meet on its own. It may not get materially worse at predicting the held-out answers. It may not get worse at returning well-formed output that downstream systems can actually parse. And it has to pass an anti-collapse probe: sample it repeatedly and count how many genuinely different answers come back, against a floor fixed in advance, because the classic failure of an over-trained specialist is not being wrong - it is answering many different questions with the same canned response. A candidate that has started doing that is refused even if its other numbers look fine. Pass, and it is adopted. Fail, and adoption is blocked and stays blocked until a person looks at the report. The system does not retry a failed evaluation until it happens to pass.

That is not a description of an intended process. A fully automatic round has now run the whole way through - selected, assembled, trained, evaluated against the full gate rather than the abbreviated trial one, and adopted. On two hundred held-out examples the trained candidate came back with roughly a third less error than the baseline it was measured against, on a gate that only asked it not to get worse. It returned valid structured output on every single case that called for it, and nine in ten of its answers across the probe set were genuinely distinct, against a floor of one in two. The loop closes.

Worth separating two things that are easy to blur, because the honest claim is the narrower one: a round completing and a candidate being adopted happen on the training machine. Putting that adopted model in front of production traffic is the separate, deliberately gated step described next, and it is the re-pinning of identity records - not the training result - that decides when callers will accept a new version.

Then there is the step most model projects leave as a manual afterthought: getting the improved model into production. That path is a defined rail - merge the new training into the base, convert and compress it, ship it to the serving machine, publish it, re-pin the identity records, and re-run the contract check. The re-pinning is the part that matters. Callers check the model that answers them against the pinned identity, so publishing a new version without re-pinning does not sneak it into production; it makes every caller reject it. The same mechanism that stops an unapproved model from serving also stops an unannounced upgrade.

03 / The vertical

The vertical is the moat

Metis Vertical Language Model Metis is shaped by company language, workflows, policies, industry signals, and market patterns, then serves agents and business workflows. Metis is a small model shaped around one enterprise vertical The goal is depth in the work that makes the enterprise different. LANGUAGECompany terms WORKFLOWSReal handoffs POLICIESAllowed moves METIS Vertical Language Model small, focused, governed AGENTSBetter calls WORKFLOWSFaster answers TEAMSShared judgment A frontier model is broad. Metis is narrow on purpose: it turns the enterprise vertical into model behavior.
Visual 1: Metis is the enterprise vertical model: narrow on purpose, useful because it is close to the work.

The vertical is not just an industry label. It is the accumulated shape of the work: how people phrase problems, which facts change the answer, which exceptions matter, which approvals are required, and which examples count as good judgment.

Metis brings that vertical into the model itself. Instead of asking a general model to borrow context at the moment of use, Metis carries the context as its native operating domain.

04 / Vertical examples

How Metis can be used across industries

Metis is useful wherever a company needs repeatable judgment inside a specialized field. The common pattern is simple: the domain has its own language, signals, evidence, risks, and decision standards, and the enterprise wants those standards to compound inside its own boundary.

That makes Metis a model for Enterprise Intelligence across finance, automotive, supply chain, ESG, telco, media and entertainment, deeptech, agritech, spacetech, and any vertical where the firm wins by seeing patterns earlier and judging them better.

VerticalHow Metis can be used
FinanceTrack market signals, compare investment theses, monitor portfolio risks, draft diligence briefs, and preserve the firm's house view across analysts and agents.
AutomotiveReason across platforms, suppliers, battery ecosystems, charging networks, recalls, policy shifts, and changing consumer demand.
Supply chainWatch supplier risk, disruption signals, inventory exceptions, contract terms, logistics bottlenecks, and sourcing alternatives.
ESGConnect evidence to claims, compare reporting standards, surface controversies, monitor assurance readiness, and keep sustainability language consistent.
TelcoUnderstand network investment, spectrum policy, enterprise offers, usage patterns, churn signals, and infrastructure tradeoffs.
Media and entertainmentMap rights, formats, audience signals, production slates, licensing windows, creator ecosystems, and brand safety context.
DeeptechTrack papers, patents, labs, grants, prototypes, commercialization signals, and the gap between scientific promise and market readiness.
AgritechReason across crops, inputs, yield risk, weather patterns, farm operations, distribution, regulation, and adoption signals.
SpacetechMonitor launch capacity, satellite markets, spectrum, defense demand, earth-observation use cases, and supplier constraints.
Metis enterprise intelligence verticals Metis turns enterprise signals into governed vertical intelligence for finance, automotive, supply chain, ESG, telco, media and entertainment, deeptech, agritech, and spacetech. Metis turns domain signals into Enterprise Intelligence The model carries the vocabulary, evidence, and judgment standards of a vertical. SIGNALSMarket changes EVIDENCETrusted facts JUDGMENTHouse view METIS Enterprise Intelligence VLM vertical, governed, current FINANCESignals AUTOPlatforms SUPPLYRisk ESGProof TELCONetworks Other verticals include media and entertainment, deeptech, agritech, and spacetech.
Visual 2: Metis turns domain signals into enterprise-specific intelligence across verticals.
05 / The economics

Why small can be better

Enterprises do not always need the most powerful model for the most repeated intelligence work. They need the right model at the right cost, with the right behavior, available wherever the workflow needs it.

A smaller vertical model can be faster, cheaper, easier to evaluate, easier to govern, and easier to place close to the business process. It can also reduce the amount of private context that has to be sent to a broad outside model for every repeated task.

The point is not that small models are universally better. The point is that a focused model can be better for a focused job.

It is worth putting numbers on that, because “small” is doing a lot of work in the paragraph above. Our own lane specialises an open seven-billion-parameter model, and a full round - assemble, train, evaluate - completes overnight on a single desktop-class GPU. That is what makes a daily cadence affordable at all, and it is the difference between a model a firm owns and a model a firm rents.

06 / Fit

What belongs in Metis

QuestionWhy it matters
Is the task repeated?Metis is strongest when many people or agents need the same kind of judgment often.
Is the domain bounded?The clearer the vertical, the easier it is to make the model useful and measurable.
Does language matter?Company terms, acronyms, market categories, and policy language are exactly where a VLM helps.
Does cost matter?Repeated calls can become expensive when every answer goes to a large general model.
Does governance matter?A focused model is easier to test, monitor, and explain inside one domain - every answered call appends a digest-stamped route receipt naming who asked, which model was requested, and which model actually served the response. Promoting a new version is gated on a complete set of evidence, with no override path. The receipt that results is cryptographically signed, and the serving side keeps re-checking that signature while the model is in service - never trusting an earlier verdict past the signature's own expiry. When it cannot be proven, the answer is labelled unverified rather than treated as promoted.
Does speed matter?Smaller specialists can sit closer to workflows that need quick answers.

Metis is not for every question. It is for the questions a company asks so often, and so specifically, that a specialist model becomes an operating advantage.

Two smaller guards are worth naming, because they are the ones that are easiest to get wrong. The model being trained is never allowed to grade its own candidate material: the check that asks “is this reference the student?” is deliberately broad and fails closed, so a reference nobody can resolve counts as the student rather than as an independent judge. And the teacher that produces training candidates is a shared service standing behind the process, never one copy per agent, and it is given no authority to promote a model or to change what production reaches for.

That guard turned out to need widening, and the reason is worth understanding, because it is a trap any enterprise fine-tuning a model can fall into. Metis is not built from nothing; it is a layer trained on top of an existing base model. Judging which material Metis should learn from using that base is not independent at all - it lets the model’s own inherited instincts decide its curriculum, and the material a base model finds confusing is precisely the material the specialised model most needs to see. One round had already discarded tens of thousands of training rows to that mistake before it was caught. The selection judge is now a different model entirely, and a configuration that names the student, or the family the student was built on, is refused rather than run.

The same principle closes a second, quieter door. Where the model’s live answers are used to build future training material, an answer the model got right is not fed back to it - only a correction written by a different model becomes training material, and a correction the model made to itself is skipped. Learning from being corrected is how a specialist improves. Learning from being agreed with is how one goes stale.

There are also, deliberately, two different answers to an unproven model, depending on when the question is asked. At the moment someone tries to deploy one, a claim of “promoted” that cannot be verified is refused outright. While a model is already serving, that same failure downgrades the record to unverified, records the blocker, and says so loudly - rather than taking the firm’s chat down. An unverifiable claim must never be honoured, and it must never become an outage either.

07 / Placement

Where Metis sits in the enterprise stack

Metis enterprise stack Frontier models provide breadth, Metis provides vertical expertise, and enterprise agents use both for daily work. The enterprise model stack needs breadth and depth Frontier models remain useful. Metis gives the enterprise a focused model for repeatable vertical work. FRONTIER MODEL Broad generalist open-world reasoningcreative stretchunusual questions handoff METIS Vertical specialist enterprise languagetrusted domainrepeatable work serves DAILY WORK Agents assistantsworkflowsteams Use the generalist for breadth. Use Metis when the work needs enterprise depth.
Visual 3: Metis complements frontier models by specializing in repeatable enterprise work.

Metis does not make frontier models obsolete. It changes what a company reaches for first, and how often it has to reach past that. The broad model is still the right tool for open-ended reasoning, creative work, and the genuinely unusual question. Metis handles the repeatable vertical work where context, vocabulary, rules, and examples matter every day.

That gives enterprise agents a better default. They can call a model that already understands the vertical, and escalate to a broader model when the task truly needs breadth.

08 / Why it matters

A company with its own vertical models

The future of enterprise AI is not one giant model answering everything. It is a model portfolio: broad generalists for breadth, vertical models for depth, and governed memory that lets judgment compound.

Metis is the vertical model in that portfolio. It is small by design, close to the business by design, shaped around the work by design, and owned by the enterprise it serves.

That is how a company moves from impressive demos to durable Enterprise Intelligence. Not by asking one general model to pretend it knows every market, asset, supplier, regulation, and technical domain, but by giving each important vertical a model that understands the work inside it.

09 / Under the hood

How Metis is actually trained

Most model announcements stop at the claim. This section is the part that usually gets skipped: where the material comes from, what is done to it before anything learns from it, and what has to be true before a new version is allowed anywhere near your teams. It is deliberately semi-technical. The shape matters more than the plumbing.

Two machines, on purpose

The enterprise’s governed stores stay where they already are, on the serving side: the library of what the firm knows, with the provenance of each entry attached; the vector index built over it; and the operational record of the work itself. A scheduled extraction reads that material under one hard filter - rows belonging to this customer, plus material that belongs to no customer at all - and assembles a dated round.

That round, and only that round, crosses to a separate training machine. The training machine holds no database credential and has no route back into the stores. It receives a package, not a connection.

On our own installation that separation is not a rule the training host is trusted to keep. It sits outside the serving side’s private network and cannot address the stores or the embedding service at all, so the extraction runs on the serving side and the finished artifact is carried across. The training machine could not reach back if it tried.

De-identification runs one way

It is worth being exact about what crosses, because the two obvious guesses - a harmless trickle of statistics, or a wholesale copy of your files - are both wrong. The training machine is not fed a stream of metrics: it reads the real material, the actual documents and decisions your teams produced, because that is the only thing you can learn a house style and a set of standards from. What it does not get is the ability to resolve who anyone is.

Two passes see to that. The first is mechanical: concrete identifiers in the text - addresses, phone numbers, hostnames, file paths, and a named list of internal identities - are replaced with stable placeholders. The substitution is consistent for the whole round, so the same real value becomes the same placeholder everywhere and the prose stays coherent enough to learn from. The map that turns those placeholders back into real values is kept on the serving side and never travels with the round, so the training machine cannot reverse it. It was never given the key.

The second pass is a judgment rather than a pattern, because the risky cases are the ones no pattern catches: a person identifiable from context, or a named counterparty attached to specific commercial terms. A separate model reads each candidate and decides. A row that fails is dropped rather than patched, and the lane can be configured to fail closed, so an unavailable judge discards the row instead of waving it through. That second pass is a filter, not a proof, and it is the honest reason the claim here is "your identifiers stay behind" rather than any absolute about names.

Nothing is admitted by default

Material does not arrive trusted and then get filtered. It arrives quarantined, and something has to certify it out of that state before anything is allowed to learn from it. On the current staged intake that is not a figure of speech: every one of the 17,359 rows is sitting in quarantine, and not one has been certified.

Certifying a row means producing the things a provenance claim actually rests on - a signing identity, the policy it was certified under, per-record provenance, secure-ingest evidence for anything from outside, and a live data-rights receipt that has not expired. Missing any of it fails closed, and the report names exactly what each row lacked rather than dropping it quietly.

Underneath sits the rule that makes the rest of it mean anything: a producer cannot certify its own material. Whatever assembled a row is never what decides that row may be learned from.

Three refusals in that gate are worth naming, because a reader would otherwise assume they are policy rather than code. A certification has to agree with its own receipt field for field - the decision, the content, the policy it was issued under, who issued it, and any blockers recorded against it - so an altered certificate fails rather than passes. A row whose usage rights are not currently live and bound to a receipt is refused outright; expired terms do not soften into a warning. And material that came out of a model is barred from the curriculum by name, so a specialist cannot quietly be fed its own output and drift away from the world it was built to learn.

The gap: what your own store does not know

No internal store is complete, and the holes are rarely where anyone expects. A vertical moves, a competitor does something new, a rule changes, and the firm’s own record is thinnest exactly where it has not been looking. Training only on what you already have will faithfully reproduce that blind spot.

So the gaps are measured rather than assumed. They are read off the model’s own behaviour in service: the questions it deferred, the ones it handed off to a broader model, the ones it could not answer from the material it had. Saturation and load are deliberately not counted, because a busy lane is an operations fact, not evidence that the model lacked knowledge.

A confirmed gap becomes a bounded mission for a Scout: go and acquire material on precisely this subject. What the Scout returns is evidence with its provenance attached, and it re-enters through the same admission gate as everything else - the same ownership checks, the same de-identification, the same certification. A Scout can put material forward. It cannot wave its own material through, and it cannot write to the model. The output of that leg is a request for a certified addition, never a quiet edit of the live corpus.

Where Metis training material comes from Governed enterprise stores on the serving side are extracted and de-identified into a dated round that is shipped to a separate training machine. Where the enterprise store has a gap, a Scout is dispatched on a bounded mission and returns evidence through the same admission gate. The key that reverses de-identification never leaves the serving side. Your content crosses. Your identifiers stay behind. The training machine reads the real material. What it cannot do is resolve who anyone is. SERVING SIDE · THE ENTERPRISE’S OWN STORES The governed library provenance on every entry The vector index over it how entries relate The operational record of the work itself ADMISSION GATE Select, then de-identify this customer only one-way TRAINING MACHINE A dated round the real text, with every identifier swapped out no credential, no key WHAT THE STORE DOES NOT HAVE A measured gap not a guess: read from where the model deferred A Scout is sent one bounded mission, for exactly that subject Evidence, with its provenance A Scout fills the hole. It does not decide what the model learns: its findings re-enter through the same gate.
Visual 4: The enterprise stores never leave the serving side. What crosses to the training machine is a dated, de-identified round, and a Scout’s findings re-enter through the same gate.

The line, and what stops it

On the training machine the round moves through five ordered phases - select, assemble, train, evaluate, adopt - and each one writes a receipt when it finishes. The next pass reads those receipts before doing anything, which is why an interrupted round resumes where it stopped rather than starting over, and why a round that has already trained is never quietly trained a second time.

The training line, and what stops it Five ordered phases — select, assemble, train, evaluate, adopt — each writing a receipt when it finishes. The next pass reads those receipts before acting, so an interrupted round resumes where it stopped and a round that has already trained is not trained again. The held-out exam is separated before training, and putting an adopted model into service is a decision of its own. The line, and what stops it at each point. Every phase writes a receipt. The next pass reads them before it acts, so a round resumes instead of restarting. Select this customer only, plus unowned material receipt Assemble de-duplicate, cap, carve out the exam receipt Train custody rebuilt and reconciled first receipt Evaluate three gates, against what is in service receipt Adopt only on a pass; a failure waits receipt Interrupted during Train The next pass reads the receipts, sees Select and Assemble already finished, and resumes at Train. It does not start over. Run again after Adopt The receipts already show a completed train for this round, so it declines rather than quietly training a second time. What stops the build The exam is separated before training, so it cannot appear in the lessons. Serving is still a separate decision An adopted model reaches production only when its identity is re-pinned. Without that, callers refuse it.
Visual 5: The training line. Every phase writes a receipt, so an interrupted round resumes where it stopped and a round that has already trained is not trained again — and the held-out exam is separated before training, while putting an adopted model into service stays a decision of its own.

Those receipts are not a log. Each one fingerprints what its phase produced - the assembled dataset, the trained adapter, the evaluation - and carries the fingerprint of the receipt before it, so a finished round is a chain that can be checked end to end rather than an account of what someone says happened.

On our own installation the lane runs autonomously, once a day.

Two things end a build rather than degrade it. The held-out examination is carved out of the de-duplicated pool before training, so it cannot turn up in the lessons and flatter the result. It is sealed rather than merely set aside: the panel is signed, and getting at it spends a single-use receipt, so a candidate cannot be quietly re-run against the exam until it happens to pass. And a final audit reads every shipped row: if a concrete identifier survived the scrub, the build stops there rather than shipping a dataset that carries it.

The trainer also checks its own input rather than trusting the package. Every row carries the terms it was admitted under, including when those terms expire. Before a single training step runs, the trainer reconciles each row against its own receipt and against the dataset’s declared minimum, and refuses the whole set if they disagree.

Who is allowed to mark the work

The judge that decides which material enters the curriculum is never the model being trained, and never the base model that model was built on top of. That second half is the one that is easy to get wrong: a specialised model is usually a layer trained over a general one, and judging its curriculum with the general model underneath lets the student’s own inherited instincts pick its lessons. The material a base model finds confusing is exactly the material the specialist most needs to see.

Three gates, then a separate decision

A trained candidate is measured against whatever is already in service, on the held-out set - which in a daily lane means yesterday’s model. Two gates are comparisons: it may not get materially worse at predicting the held-out answers, and it may not get worse at returning output that downstream systems can parse. The third is an absolute standard - sample it repeatedly and count how many genuinely different answers come back, against a floor fixed in advance - because the classic failure of an over-trained specialist is not being wrong, it is answering everything with the same canned response. A pass adopts. A failure blocks adoption and stays blocked until a person reads the report.

Who does the marking, and why it is not one of them

Those three gates are arithmetic. Loss against the held-out answers, whether the output still parses, how many genuinely different answers come back — each one is counted, and no model is asked for an opinion. That is deliberate: the cheapest way to fool yourself is to let something with a stake in the answer decide whether the answer is good.

The richer measurement sits alongside them, and it does need judgement. The candidate is put head to head against the model already in service on held-out prompts, and the two answers are compared. The judge is a frontier model — a large general model, deliberately more capable than either contestant at the narrow task of telling a good answer from a poor one.

What matters is what it is not. It is not either contestant. It is not the student, and it is not the model the student was built on top of — that second exclusion is the one that is easy to miss, because a specialist is a layer trained over a general model, and letting that general model mark the work lets the student’s own inherited instincts decide what counts as good. And because a judge shown the same pair twice can contradict itself, every pair is put to it in both orders. Only a consistent winner is scored. When it disagrees with itself, that is recorded as a tie rather than resolved in someone’s favour.

Who marks the work The candidate and the model already in service are compared by a frontier model acting as judge. The judge is neither contestant and shares no lineage with them. Every pair is put to it twice, once in each order; only a consistent winner counts, and disagreement is recorded as a tie. The three deterministic gates involve no model at all. The judge is a frontier model with no stake in the result. Neither contestant marks its own work — and neither does anything built on the same lineage. THE TWO CONTESTANTS the candidate this round’s adapter already in service in a daily lane, yesterday’s model a frontier model larger than either contestant not the student not the model the student was built on it is not competing every pair is put to it twice, once in each order WHAT COUNTS same winner both ways the pair is scored it contradicts itself recorded as a tie, not a win The three gates are arithmetic. No model’s opinion enters them. Loss against the held-out answers · whether the output still parses · how many genuinely different answers come back
Visual 6: Who marks the work. The judge is a frontier model that is neither contestant and shares no lineage with them; every pair is put to it twice, and a judge that contradicts itself scores a tie.

Adoption is still not deployment. Putting an adopted model in front of your teams is its own gated step, and the part that carries the weight is re-pinning the model’s identity records. Callers check the model that answered them against what is pinned, so publishing a new version without re-pinning does not sneak it into production. It makes every caller reject it. The same mechanism that stops an unapproved model from serving also stops an unannounced upgrade.

What a pilot settles

The factory is the product; your model is the output

Everything above runs today on our own work. What a pilot establishes is the part that can only be done with you: which vertical, which of your stores feed it, where your gaps actually are, and what "good" looks like on the work your teams do every day.

Ownership and licensing. Foundation-AI and Foundation-LifeStyle, together with all intellectual property rights subsisting in them, are the sole and exclusive property of MediaGlyphics GK. Ibex is an authorized licensor of these technologies for forward deployed engineering (FDE) engagements. © 2026 MediaGlyphics GK. All rights reserved.