Levanto's Sage Was Ranking Tripcite's Audit Queries in Production 50 Days Before TypeSafe AI Came Out of Stealth With $40 Million for the Same Idea
A decision model takes content and a question and returns a typed answer with a calibrated confidence, not prose, in a single forward pass. TypeSafe AI announced one on 15 September with a $40 million seed round led by DCVC. Levanto Labs published its own, Sage, to npm on 13 July, and Tripcite has been running it on every self-serve audit since 27 July: Claude writes 30 candidate questions, Sage ranks them, and the audit measures the top 14. In the test that earned it the job, Sage put all 14 real queries above all five planted bad ones.

On 15 September TypeSafe AI, a San Francisco lab founded by former OpenAI researcher Diogo Almeida with Erik Gafni and Sasha Sheng, came out of stealth with $40 million in seed funding led by DCVC and a model called Jev that returns typed decisions with calibrated probabilities instead of text, in under 100 milliseconds. Forbes reported a $200 million valuation. Sixty-four days earlier, on 13 July at 08:12 UTC, a company with no announced funding had published a client for the same kind of model to npm. Levanto Labs, founded by chief executive Marco De Rossi with Christopher Kocurek, previously marketing director for Linea at ConsenSys, calls its model Sage. Its GitHub organisation was created on 8 July. Two weeks after the npm publish, Sage was in production at tripcite.com, the AI-search visibility tool built by the same team as this desk, where it has ranked the questions behind every self-serve audit since 27 July.
The category both companies are selling is new enough that each has had to name it. TypeSafe calls Jev a System One model, after Kahneman's fast intuitive thinking. Levanto calls Sage a decision model, "for workflows where software needs to decide, not write". The mechanism is the same and it is the interesting part.
A model that answers in one pass
A chat model writes its answer one token at a time, and a program that wants a yes or a no has to parse it out of prose, retry when the format drifts, and take the answer on faith because a language model's stated confidence is a sentence, not a measurement. A decision model is asked a closed question about a piece of content and produces a probability distribution over the possible answers in a single forward pass. There is nothing to parse. The answer field is a value the code can branch on, and the shape of the distribution is a number the code can threshold: concentrated on one option means the model is sure, spread across several means it is not.
That second number is the product. Levanto's line is that "without a calibrated signal, you trust everything or nothing", and its pitch is a loop in which software acts above a confidence threshold and escalates to a human below it, in 60 to 200 milliseconds. TypeSafe describes the same loop, documents the same distribution-shape definition of confidence, and recommends the same three bands: above 0.9 act, 0.5 to 0.9 confirm, below 0.5 hand to a person. TypeSafe says Jev is trained with a method it calls Reinforcement Learning for Calibrated Decisions so that the probabilities are honest rather than merely present. Levanto's homepage describes Sage as "a 300B model, responding in 100ms". TypeSafe has published no weights.
| Sage, Levanto Labs | Jev, TypeSafe AI | |
|---|---|---|
| First public code | 13 Jul 2026, npm 'levanto' 0.1.0 | 15 Sep 2026, npm '@typesafe-ai/sdk' |
| Company out of stealth | GitHub org 8 Jul 2026 | 15 Sep 2026, $40m seed led by DCVC |
| Decision kinds | Yes/no, choice, scale, sort, tags | Choice, score, true/false |
| Largest list | 120 items to sort, 120 options to choose | 255 options to choose |
| Latency, vendor-stated | 60 to 200 ms | Under 100 ms; 70 to 500 ms end to end |
| Confidence | 0 to 1 per decision, null when unsure; list-level for sort | 0 to 1 from the distribution shape; none on true/false |
| Model, vendor-stated | "A 300B model" | New architecture, parallel sampler, RLCD; no weights published |
| Pricing | Per decision: free developer tier, $14/mo for 30,000 | $0.042 per million input tokens, output free |
| Named integrations | Tripcite (27 Jul), Deepsec (14 Sep) | Early access, waitlist |
Where the two differ is in what a caller can ask. Jev exposes three primitives: a choice from up to 255 options, a score against ordered levels, and a true or false. Sage exposes five, and the two extras are the ones Tripcite ended up depending on: a fixed five-level rubric it calls scale, and sort, which takes a list of between 2 and 120 items and returns them in order with a single list-level confidence for the whole ordering. Sage's current version also returns null when it is not sure, rather than a low number, which turns the human hand-off from a threshold the developer has to pick into a case the API forces them to handle.
What Tripcite needed it for
A Tripcite audit reads a company's website, writes the questions a prospective buyer would put to an AI assistant, asks five engines each of them (Perplexity, Gemini, ChatGPT search, Claude and Google AI Overviews) and records whether the business is cited in the answer. The free audit is capped at 14 measured queries. The question set is the highest-leverage input in the run: each kept query costs about five engine calls, and a badly formed one drags the headline mention rate down for a reason that has nothing to do with the client's visibility. Claude writes good questions most of the time. The problem was that nothing checked them.
The first design was a filter. On 27 July Tripcite ran about 130 Sage decisions on a trial account against a fixture of 24 labelled queries from two unrelated businesses, a Spanish sports-travel operator and a US developer-tools API, with the bad examples built from the failure modes that actually occur: keyword fragments, over-broad prompts, off-topic and unwinnable brand questions. Three question shapes were scored head to head, and the metric was AUC, the probability that a randomly chosen good query outscores a randomly chosen bad one, where 1.0 is perfect and 0.5 is a coin flip.
The five-level rubric scored 0.965. The yes/no question the brief had proposed, "would a real buyer type this?", scored 0.903. A count of flaw tags scored 0.847, and two of the five tags turned out to be pure noise at 0.514: "keyword, not conversational" fired on the query "how do I get a calibrated confidence score back from an LLM so my code can decide when to escalate", which is the opposite of a keyword fragment. A gate that used the tags as vetoes sent 100 percent of the good queries back for regeneration. With those two tags removed and the rubric as the only judge, the fixture came out 23 of 24 correct.
Then the gate met real output and did nothing. On the 14 questions Tripcite's live pipeline generated for levanto.ai, 14 were kept, none regenerated, none dropped, every one scoring between 3.95 and 4.00 on the four-point rubric with confidence between 0.91 and 0.98. The generation prompt already demands full conversational questions grounded in the site's own text, so there was almost nothing to filter, and a score that saturates at the top cannot say which of fourteen good candidates is better than another. Choosing between good candidates is the whole job.
Sort cannot saturate
The third shape fixed exactly that. The 14 real queries were spiked with five planted bad ones, "LLM API", "AI", "best AI tool", "how do I train a large language model from scratch" and "what is the best AI company in the world", shuffled into alphabetical order so that position could not explain the result, and sent as one list with one instruction: rank these by how well each would test whether an AI assistant recommends this business.
All 14 real queries came back above all five plants, ranks 15 to 19, the perfect ordering, at a list confidence of 0.718. The order inside the real block was sensible too. Top: "we need an AI that can make routing decisions in under 200ms for our agentic workflow". Bottom of the real block: "Levanto reviews for agentic AI workflows", the thinnest and most keyword-like of the set. Sort is comparative, so two candidates cannot both be a 4.00; one of them has to go above the other. That ordering is a capability Tripcite did not have at any price before, because a chat model asked to rank 19 strings returns prose about them.
How it runs now
The live wiring is one call at one point in the flow. Claude is asked for 30 candidates instead of 14, at marginal cost since it is the same request. Sage sorts the pool once and the top 14 arrive in the wizard pre-selected, credited on the page as "Ranked by Sage · Levanto Labs". The other 16 ride along unticked so a user can promote a spare instead of retyping. Bilingual audits generate 18 candidates per language and rank Spanish and English separately, keeping eight of each and interleaving them, because ranking a Spanish query against an English one measures the translation rather than the query. Measured on odisea-tours.com the day it shipped, the 30 candidates generate in about 20 seconds and the ranking adds about two, with sort confidence holding at 0.705 to 0.707 even when the pool was pushed to 50 items. If Sage is unreachable the queries keep their original order and the attribution line is not shown.
Sort is billed per item, not per call. On the live account a 30-candidate sort billed 27,964 input tokens, a 50-item sort 47,010 and a 19-item sort 11,810, roughly 600 to 1,000 tokens per item, which means the business context sent with the list is paid for once per candidate. Tripcite cut that context from 1,200 characters to 400, removing about 6,000 billed tokens per audit. Before the cut, ranking cost $0.084 per audit under the token pricing in force in July, against about $0.49 in engine calls for a free audit. Levanto has since moved to per-decision plans, from $14 a month for 30,000 decisions with a free developer tier, read on 21 September. TypeSafe prices Jev at $0.042 per million input tokens with output free.
Two things are untested. The ranking instruction is English while Spanish candidates are not, and the script written to check whether that matters could not run because the trial credit was exhausted. And the fixture is 24 queries and one live domain, with the bad examples written by Tripcite rather than sampled from real generation failures, because real failures turned out to be rare. The AUC figures read as clearly works, not as precise numbers.
Who was first
The public record is unambiguous on the order. Levanto's GitHub organisation dates from 8 July, its JavaScript and Python clients from 13 July, Tripcite's production integration from 27 July, and on 14 September a security scanner called Deepsec merged a Sage triage gate that discards obviously benign regex matches before an expensive coding-agent investigation, its author's pull request records. TypeSafe's release, its Jev early-access programme and its npm SDK are all dated 15 September. What TypeSafe has that Levanto does not is $40 million, a founder its release credits as a co-inventor of the training method behind ChatGPT, and vendor-run benchmarks claiming Jev is 193.6 times faster and 244.6 times cheaper than a chat model on the same task. What Levanto has is 64 days, two extra decision kinds, and paying integrations that already depend on them.
- Tripcite, self-serve audit (the query step carries the Sage ranking)
- Levanto Labs, Sage
- Levanto docs, sort decisions
- Levanto docs, limits
- Levanto docs, pricing
- npm registry, levanto (published 13 July 2026)
- GitHub, levantolabs (organisation created 8 July 2026)
- GitHub, Deepsec pull request 2: Levanto Sage triage provider (merged 14 September 2026)
- TypeSafe AI
- TypeSafe docs, confidence
- SiliconANGLE, TypeSafe AI exits stealth with $40M (16 September 2026)
- DCVC, TypeSafe emerges from stealth
- MarkTechPost, TypeSafe AI releases Jev (19 September 2026)
- npm registry, @typesafe-ai/sdk (published 15 September 2026)
We report facts in our own words and link to the reporting we drew them from. We do not reproduce a source's prose, headline or images. Nothing here is investment advice.