ATōMIC is a local-first reasoning layer for small models. It measures what the model actually knows, abstains when it's unsure, and shows its work — built for teams where a wrong answer costs more than no answer.
The models keep improving and still confabulate. In regulated work that gap isn't theoretical — it's in the case law, the clinical studies, and the examiner's checklist.
The research already named the fix: reward calibrated abstention, not guessing. EU AI Act transparency duties went live Aug 2026; high-risk lands Dec 2027 — so the teams that start now are the ones who'll be ready.
It sits between your model and its answer. Nothing here is a claim the benchmark can't back.
The reliability vendors are strong, and most now ship on-prem and VPC. But they all work the same way: a verdict or a score on a finished answer. None carry a claim-level confidence trail, separate "the document says X" from "I infer X," or remember whether a fact is still true across sessions. That combination, during reasoning, is the seat nobody's in.
The fair question from every technical buyer. Native calibration and refusal are improving — but they stop short of what an audit needs.
We're pre-product with zero shipped customers, so we don't ask for trust. The Calibration Bench ships open with v0.1 — you run it yourself, and we label every number by who measured it.
Small models — 8B, shrinking toward ~300M — mean the reasoning happens where your data already lives. No BAA chain to a frontier vendor, no data-processing addendum, no third-country cloud.
Every scenario is grounded in documented harm — not a customer story we don't have yet. This is what ATōMIC is for.
The developer entry point into epistemic reasoning: open core, runs locally, quickstart in minutes. Shipping with v0.1 — with an honest note on what's in the first release and what's next.
Capped at a few partners — ideally one in legal, one in medical, one in financial. Real scarcity, said honestly. The exchange:
Our only social proof is the truth: real people, real research, in the open. [Draft — swap in the real team, photos, and bios.]
Join the drop list for the open benchmark and toolkit, or apply to build the first deployment in your vertical.
Draft — form is a mockup; wire it to your CRM and verify the inbox before launch.