AI that knows when it doesn't know.

ATōMIC is a local-first reasoning layer for small models. It measures what the model actually knows, abstains when it's unsure, and shows its work — built for teams where a wrong answer costs more than no answer.

Pre-release. Every benchmark is open and reproducible. Design partnerships open now.
atomic · on-device · 8Breasoning
What's the indemnity cap in the Meridian MSA?
The cap is $2.0M, per §11.3 of the executed MSA.
confidence0.94
trace → matched clause MSA §11.3 · exact figure · 1 source · no inference
Does that cap apply to the data-breach carve-out?
I'm not sure — the carve-out isn't in the documents I've been given.
confidence0.41
what would change this → add the Security Addendum or Schedule C and I'll re-answer with a citation.

Confidently wrong is now a documented, sanctioned risk.

The models keep improving and still confabulate. In regulated work that gap isn't theoretical — it's in the case law, the clinical studies, and the examiner's checklist.

~1,490
court decisions worldwide involving AI-fabricated content — with the first attorney bar suspensions handed down in 2026.
Charlotin tracker · GC AI · May 2026
31%
of AI-generated clinical notes contain hallucinations. The physician who signs carries the exposure.
peer-reviewed · 2026
1.8–3.7%
hallucination rate for the best models on a summarize-only-from-facts task. Fatal in a filing or a credit memo.
Vectara HHEM-2.3 · 11 May 2026
SR 26-2
the first US model-risk revision in 15 years leaves banks needing an audit-ready reasoning trail for AI outputs.
OCC / Fed / FDIC · Apr 2026

The research already named the fix: reward calibrated abstention, not guessing. EU AI Act transparency duties went live Aug 2026; high-risk lands Dec 2027 — so the teams that start now are the ones who'll be ready.

A layer that measures what the model knows — and refuses to bluff.

It sits between your model and its answer. Nothing here is a claim the benchmark can't back.

01 · MEMORY
It remembers — and knows if it's still true.
Long-term memory across sessions, carrying epistemic status. Facts age, get re-verified, or get downgraded, so a stale fact never resurfaces as a truth.
02 · UNCERTAINTY
Every claim carries a confidence it has to earn.
It tells you how sure it is, and separates "the document says X" from "I infer X." A number grounds to a source or it doesn't ship.
03 · ABSTENTION
Below the bar, it says "I'm not sure."
Calibrated abstention — right answers kept, wrong answers refused — instead of filling silence with fiction. And it tells you what evidence would change the answer.
04 · TRACE
Every answer ships with its work.
Sources, steps, and confidence attached to the output — the audit trail a regulator, a reviewer, or opposing counsel can walk through line by line.

Everyone else checks the answer after. On someone else's cloud.

A reasoning layer that manages uncertainty while it reasons — and runs where your model already runs — is an empty quadrant.

The hallucination-detection vendors are cloud platforms that judge outputs after the model speaks, adding latency and still false-positiving. The small-model labs don't do reliability. The sovereign players don't publish mechanism. Calibrated reasoning, local-first, is the seat nobody's in.

Proof you can run — not logos we don't have.

We're pre-product with zero shipped customers, so we don't ask for trust. We publish a benchmark you run yourself, and we label every number by who measured it.

The Calibration Benchmetric
Confident-error rate — wrong, and sure of ittarget ↓
Calibrated-abstention — kept vs. refused, correctlythe number
Over-abstention — good answers wrongly refusedscored, not hidden
On-device runtime overheadmeasured per run
Run it on your own datano egress
# pull the harness $ git clone github.com/nola-ai/calibration-bench # score any model, on your machine $ atomic bench --model ./your-model \ --data ./your-eval-set --local # → confident-error · calibrated-abstention # over-abstention · latency — reproducible
Why abstention is scored both ways: a model that just says "I don't know" a lot is useless. The bench rewards calibrated abstention — keeping the answers it should, refusing only what it can't ground. Every NOLA figure is marked self-reported until an outside party reproduces it, then verified. You'll always know which.

Local-first isn't an option. It's the point.

Small models — 8B, shrinking toward ~300M — mean the reasoning happens where your data already lives. No BAA chain to a frontier vendor, no data-processing addendum, no third-country cloud.

On-device / on-prem
Runs on your hardware. PHI, privilege, and PII never leave the building.
default
Air-gapped
Fully disconnected environments. No network egress, verifiable.
highest assurance
Your VPC
Your cloud account, your keys, your perimeter — when you want ops handled but control kept.
managed

Built for the work where a wrong answer has a cost.

Every scenario is grounded in documented harm — not a customer story we don't have yet. This is what ATōMIC is for.

Touch it today. Or build it with us.

ATōMIC ToolKit community edition

The developer entry point into epistemic reasoning. Open core, runs locally, quickstart in under five minutes — with an honest note on what's in this release and what's next.

$ pip install atomic-toolkit
Read the docs

Design partners 3–5 seats

Capped at a few partners — ideally one in legal, one in medical, one in financial. Real scarcity, said honestly. The exchange:

  • You bring a real workload, and data access inside your own environment
  • An hour every two weeks, plus permission to publish anonymized results
  • You get a working deployment on your hardware + direct influence on the roadmap
  • Founding-partner pricing, locked — and co-authorship on results if you want it
Apply to partner

An AI stack a regulator can walk through, end to end.

ATōMIC
the reasoning layer — knows what it knows
Data Room Protocol
how the record is kept — structured, sourced, sovereign
P3AK vault
where it lives — encrypted, portable, yours

ATōMIC reasons over your data, and how that data is organized matters. NOLA works with P3AK, the private-brain data layer from mpressed: encrypted, portable vaults you own outright, organized under the open Data Room Protocol so regulated documents stay structured, sourced, and sovereign on your own hardware. Each layer works alone. Together, they let an auditor trace any answer all the way back to its source.

Deep-tech, built in New Orleans.

Our only social proof is the truth: real people, real research, in the open. [Draft — swap in the real team, photos, and bios.]

Founder / CEO
name · bio
Chief Scientist
name · bio
Head of Research
name · bio
Engineering
name · bio

The benchmark is open. The seats are few.

Run the Calibration Bench on your own data, or apply to build the first deployment in your vertical.

DRAFT REDESIGN · concept by mpressed / P3AK for NOLA AI · copy grounded in dated, sourced research (Aug 2026) · benchmark figures are placeholders pending real runs · not published to nola-ai.com