AI that knows when it doesn't know.

ATōMIC is a local-first reasoning layer for small models. It measures what the model actually knows, abstains when it's unsure, and shows its work — built for teams where a wrong answer costs more than no answer.

Pre-release. The open benchmark and toolkit ship with v0.1 — design partnerships open now.
atomic · on-device · 8Breasoning
What's the indemnity cap in the Meridian MSA?
The cap is $2.0M, per §11.3 of the executed MSA.
confidence0.94
trace → matched clause MSA §11.3 · exact figure · 1 source · no inference
Does that cap apply to the data-breach carve-out?
I'm not sure — the carve-out isn't in the documents I've been given.
confidence0.41
what would change this → add the Security Addendum or Schedule C and I'll re-answer with a citation.
Illustrative — the target behavior of the v1 trace format. Not a live run.

Confidently wrong is now a documented, sanctioned risk.

The models keep improving and still confabulate. In regulated work that gap isn't theoretical — it's in the case law, the clinical studies, and the examiner's checklist.

~1,490
court decisions worldwide involving AI-fabricated content — with the first attorney bar suspensions handed down in 2026.
Charlotin tracker · GC AI · May 2026
~7M
medical visits ran through a transcription tool that fabricated text never spoken — ~40% of those hallucinations judged potentially harmful.
"Careless Whisper" study · CIO / ACDIS
1.8–3.7%
hallucination rate for the best models on a summarize-only-from-facts task. Fatal in a filing or a credit memo.
Vectara HHEM-2.3 · 11 May 2026
SR 26-2
rescinded SR 11-7 and left generative AI outside formal scope — so banks must still evidence controls on any AI output that sways a decision.
OCC / Fed / FDIC · Apr 2026

The research already named the fix: reward calibrated abstention, not guessing. EU AI Act transparency duties went live Aug 2026; high-risk lands Dec 2027 — so the teams that start now are the ones who'll be ready.

A layer that measures what the model knows — and refuses to bluff.

It sits between your model and its answer. Nothing here is a claim the benchmark can't back.

01 · MEMORY
It remembers — and knows if it's still true.
Long-term memory across sessions, carrying epistemic status. Facts age, get re-verified, or get downgraded, so a stale fact never resurfaces as a truth.
02 · UNCERTAINTY
Every claim carries a confidence it has to earn.
It tells you how sure it is, and separates "the document says X" from "I infer X." A number grounds to a source or it doesn't ship.
03 · ABSTENTION
Below the bar, it says "I'm not sure."
Calibrated abstention — right answers kept, wrong answers refused — instead of filling silence with fiction. And it tells you what evidence would change the answer.
04 · TRACE
Every answer ships with its work.
Sources, steps, and confidence attached to the output — the audit trail a regulator, a reviewer, or opposing counsel can walk through line by line.
Works with
ATōMIC reasons over your data, and how that data is organized matters. NOLA works with P3AK (mpressed) and the open Data Room Protocol — encrypted, portable vaults you own, so regulated records stay structured, sourced, and sovereign on your own hardware. Each piece works alone; together they make a stack an auditor can trace end to end.

Everyone else checks the answer after it's written.

A layer that manages uncertainty while it reasons — claim by claim, with memory — is an empty quadrant.

The reliability vendors are strong, and most now ship on-prem and VPC. But they all work the same way: a verdict or a score on a finished answer. None carry a claim-level confidence trail, separate "the document says X" from "I infer X," or remember whether a fact is still true across sessions. That combination, during reasoning, is the seat nobody's in.

"Won't the model just do this itself?"

The fair question from every technical buyer. Native calibration and refusal are improving — but they stop short of what an audit needs.

A logprob isn't a citation.
A model's own confidence is a number with no receipt. ATōMIC ties confidence to a claim-level trace — which source, which step, stated vs. inferred — the thing a reviewer can actually check.
Native refusal has no memory.
A fine-tune can abstain in the moment, but it can't remember that a fact was true in March and changed in the amended filing. Epistemic status that persists across sessions is a layer, not a weight.
It has to run where your data is.
The abstention-trained models the 2026 literature is producing are frontier-scale and cloud-served. ATōMIC targets a 300M-class model on your own hardware — reliability where the regulated data already lives.
And it has to be provable.
"Trust the model's judgment" is what a regulator won't accept. The Calibration Bench is the referee: the delta over native calibration is a number you run, not a claim we make.

Proof by an open benchmark — not logos we don't have.

We're pre-product with zero shipped customers, so we don't ask for trust. The Calibration Bench ships open with v0.1 — you run it yourself, and we label every number by who measured it.

The Calibration Benchmetric
Confident-error rate — wrong, and sure of ittarget ↓
Calibrated-abstention — kept vs. refused, correctlythe number
Over-abstention — good answers wrongly refusedscored, not hidden
On-device runtime overheadmeasured per run
The open harness — v0.1 previewnot yet released
# ships open with v0.1 — join the drop list below. # the planned interface: score any model on your # own machine, nothing leaves the network. $ atomic bench --model ./your-model \ --data ./your-eval-set --local # → confident-error · calibrated-abstention # over-abstention · latency — reproducible
Why abstention is scored both ways: a model that just says "I don't know" a lot is useless. The bench rewards calibrated abstention — keeping the answers it should, refusing only what it can't ground. Every NOLA figure is marked self-reported until an outside party reproduces it, then verified. You'll always know which.

Local-first isn't an option. It's the point.

Small models — 8B, shrinking toward ~300M — mean the reasoning happens where your data already lives. No BAA chain to a frontier vendor, no data-processing addendum, no third-country cloud.

On-device / on-prem
Runs on your hardware. PHI, privilege, and PII never leave the building.
default
Air-gapped
Fully disconnected environments. No network egress, verifiable.
highest assurance
Your VPC
Your cloud account, your keys, your perimeter — when you want ops handled but control kept.
managed

Built for the work where a wrong answer has a cost.

Every scenario is grounded in documented harm — not a customer story we don't have yet. This is what ATōMIC is for.

Touch it today. Or build it with us.

ATōMIC ToolKit community edition

The developer entry point into epistemic reasoning: open core, runs locally, quickstart in minutes. Shipping with v0.1 — with an honest note on what's in the first release and what's next.

$ # open install lands with v0.1
Join the drop list

Design partners 3–5 seats

Capped at a few partners — ideally one in legal, one in medical, one in financial. Real scarcity, said honestly. The exchange:

  • You bring a real workload, and data access inside your own environment
  • An hour every two weeks, plus permission to publish anonymized results
  • You get a working deployment on your hardware + direct influence on the roadmap
  • Founding-partner pricing, locked — and co-authorship on results if you want it
Apply to partner

Deep-tech, built in New Orleans.

Our only social proof is the truth: real people, real research, in the open. [Draft — swap in the real team, photos, and bios.]

[ Name ]
[ role ]
[ Name ]
[ role ]
[ Name ]
[ role ]
[ Name ]
[ role ]

Get in early.

Join the drop list for the open benchmark and toolkit, or apply to build the first deployment in your vertical.

Or email us

Draft — form is a mockup; wire it to your CRM and verify the inbox before launch.

DRAFT REDESIGN · concept by mpressed / P3AK for NOLA AI · copy grounded in dated, sourced research (Aug 2026) · benchmark figures are placeholders pending real runs · not published to nola-ai.com