We stopped building for a week to measure what we'd built. Here are the numbers.
By ATLASS OS team · August 31, 2026 · documents an event on August 30, 2026
We build ATLASS OS in checkpoints: stop building, measure what we have, make sure it's true, publish the numbers, then build the next thing. This is the checkpoint after a two-week sprint that took the platform's AI surface from 58 tools (August 7) to 184 (as of 2026-08-30, commit aa18ea6b) — every one of those counts produced by the same script, from the deployed code, on the date shown.
Where the surface stands now: 195 registered tools (196 on the wire, 118 write-capable across 72 scope groups), counted at commit 4fb20da4 on 2026-09-07 by scripts/tool-inventory.mjs --json. The numbers in this post are the ones it was published with and are left as they were.
What ATLASS OS is. A business operating system for trades and field-service companies: CRM, quotes and invoices, payments recorded, jobs and scheduling, full double-entry books with Canadian tax, Canadian payroll, and a business phone line with missed-call text-back — one system, one ledger, priced in CAD. It is founder-run; the founder's roofing company runs on it, and his two other businesses are tenants on it today. It is in open beta by reservation today: sign up at the founding form and we set your business up with you — self-serve is next. Run a real month-end on it and tell us where it isn't good enough yet. That is the point of a beta. Two honesty lines up front, because we'd rather you read them here than find them in week one: payroll and sales-tax returns are calculated and recorded for you — you still submit them to the CRA yourself and log the confirmation; and owners can export a full backup of their books any day (JSON today, CSV coming).
What "speaks fluent AI" means here. Connect the AI agent you already use — Claude, ChatGPT, any MCP client — with a token you mint inside the app. You choose the scopes; the token can only ever use what it was granted when it was made; every action lands in an append-only audit log; and nothing in ATLASS, human or AI, can move a dollar, because there is no payment rail in the product. Your agent can work your CRM, your books, and your schedule through the same audited doors a person uses. Connecting is included on every plan.
The number nobody else in this category has published — what YOUR OWN AI assistant achieves when it drives ATLASS through MCP. In our pre-registered benchmark, Claude Sonnet 5 selected the right tool or a valid first step first time 90.3% of the time, and Claude Haiku 4.5 — the cheapest tier — 66.7%, on the 184-tool surface as of 2026-08-28 (iterating 3194da63). Selection, not task completion. Your own agent, your own numbers. How it was run: 72 plain-language office asks ("who owes me money", "pay the roofing supplier's bill", "start the payroll run") written from the real tool surface, the full catalog on the table, chain-aware, first-turn tool selection, single attempt, no retries; every non-hit judged independently under a fixed rubric. Haiku understood 100% of the asks and made zero wrong-tool picks; Sonnet's comprehension was 97.2%. The lower number is the one we commit to raise — the cheapest tier is our bar. Across all 288 recorded decisions there were four genuine wrong-tool picks — all by the stronger model, all in tax questions; every other non-hit was a sensible clarifying question.
This measures tool selection — which door your agent reaches for first — not end-to-end task completion; that is a separate, more expensive test we'll publish when we run it. We say "first try, single attempt" because retries flatter every benchmark, and we name the models because "an AI" is not a number. (These are numbers for the agent you bring; they describe nothing else.)
The finding we hoped would be bigger, published anyway. Our pre-registered hypothesis was that department-scoped tokens (a token that can only see the accounts-receivable tools, say) would raise a cheap model's hit-rate. They didn't — neutral for Haiku (+1.4 points), slightly negative for Sonnet (−7), because scoping removed the cross-department lookup tools the models correctly reach for first. So the honest claim is narrower and, we think, better: least-privilege scoping costs the cheap tier nothing. Scope tokens down for security; you pay no selection penalty. The benchmark also found a product fix (shared lookup tools now ride every scoped token) — measuring finds things. That is why we measure.
QuickBooks, for contrast — facts only, each dated. Intuit publishes an MCP server on GitHub (145 tools; https://github.com/intuit/quickbooks-online-mcp-server, snapshot 2026-08-28). It is self-hosted and developer-setup-only: you register your own OAuth app, run a local process, and rotate tokens yourself; there is no hosted version. Intuit's turnkey Claude connector, per its own help article (https://quickbooks.intuit.com/learn-support/en-us/help-article/accounting-bookkeeping/use-quickbooks-connector-claude/L3YBlo6HtUSen_US, snapshot 2026-06-21) and its July 28, 2026 announcement (https://www.intuit.com/blog/news-social/quickbooks-expands-into-claude-and-chatgpt-with-new-sales-invoicing-payroll-and-lending-features/), creates, updates and sends sales transactions — invoices, estimates, payment links, customers, products; journal entries are not among its listed use cases, and the article says plainly that "the current QuickBooks connector use cases are limited to the ones above." Neither Intuit nor anyone else in the category has published a tool-selection hit-rate (searched 2026-08-30). ATLASS's server is hosted, minted in the app in about a minute, scoped per token, audited — and it reaches receivables, payables, bank, payroll and tax: 184 tools (113 of them write) across 67 scope groups at commit aa18ea6b, every posting landing in the same double-entry ledger.
This post is three days late, on purpose. Our pre-release security sweep found issues in surfaces we hadn't finished hardening. We closed every one, had each fix re-verified independently, and then published. The financial core — every door that touches money — was clean in that sweep. That is the policy here: nothing ships ambiguous. Every number on this page is machine-counted and dated, and every security claim was re-tested after its fix. SOC 2? Not yet — certification is on the roadmap; what's true today is verifiable in the product: scoped tokens, append-only audit log, no payment rail, per-token rate limits, a per-business kill switch.
What's live and what isn't. Live: everything in the first paragraph, the MCP surface, the phone line with missed-call text-back, receipt capture and field apps. Coming, and labelled that way everywhere we sell: our own in-app AI staff you talk to like employees (private beta — nothing switches on, and nothing is billed, without 30 days' notice), voice control, financial reports delivered as documents in chat, an AI onboarding guide, self-serve phone-number setup, and inviting your team from inside the app (today we add your people for you — say so in the founding form).
Pricing, in CAD, per business (locked 2026-08-30; prices as of 2026-08-30, rendered from the catalog at publish time). Core $29/mo — CRM, quotes and invoices, full double-entry books, your own AI agent connected. Pro $79/mo — adds jobs, purchase orders, progress invoicing, spatial, the field apps. Two office seats included on both; extra office seats $9 / $29. Your crew on the field apps and the phone line are not seats — a four-person shop with two in the office runs Pro at $79. Connecting your own AI agent is included on every plan and free to mint. Our own in-app AI staff: private beta — always an add-on to Pro, never bundled into a plan; pricing is set before it ships, and nothing is billed without 30 days' notice; Pro subscribers get their first month of the AI add-on free. Your trial includes $10 of AI assistant usage — about 100 conversations — on the full CRM + accounting toolset. Add-ons: Canadian payroll $15/mo + $1 per paycheque (calculated and recorded; you file); business phone $20/mo per number with 300 texts and 100 minutes a month. Annual: two months free. 30-day full trial, no card, no contracts. Your rate never rises while you stay active — list prices step up as we grow, early users keep theirs. Founding partners: $49/mo for everything in Pro, locked for as long as you stay active; a spot is secured by your first payment when billing opens, and reserving today costs nothing.
Start. Reserve a founding spot or ask for the trial at atlass-os.com/founding (self-serve signup is next — no date) · read the numbers and the methodology at app.atlass-os.com/mcp · or just ask the agent on the page; it answers from this post's facts and nothing else.
See the MCP tool surface or pricing.