A small language for AI agents
Typed model calls, budgets in money and time, tools, and concurrency.
Today’s agent frameworks are libraries. The things that matter most — how much an agent may spend, how long it may run, which tools it may touch, what shape an answer must have — live in those libraries as conventions, and nothing checks that they are followed. µNorman makes them part of the language.
The interpreter on this page is the same Rust program as the command-line one, compiled to WebAssembly. Every example below is repeatable and will run on your machine.
What a language can do that a library can’t
Six papers shaped this design. Each points at something that belongs in a language rather than in a library.
| Problem | What µNorman does |
|---|---|
| Model answers have no reliable shape | every ask names the type of answer it expects |
| Agents decide between options | answers can be choices, and programs branch on them |
| Cost and deadlines are hard limits | money and time are built in; the budget can never be overspent |
| Latency compounds | one shared clock per budget scope, with deadlines everywhere |
| A failed step shouldn’t rerun everything | failures are values, and retry wraps one step |
| Agents shouldn’t touch tools they weren’t given | tools are capabilities; a model can never create one |
Note that most of these are enforced when the program runs. See Status.
A tour via examples
1. Ask
ask takes a model, the type of answer you want, and the context the
model sees. The answer comes back as a real value of that type, or the ask fails.
Context is an ordinary list you build and pass: nothing accumulates behind your back.
(grant claude (model [in 3] [out 15] [ceiling 8000]))
(datatype Verdict
[Buy (reason (Text 400))]
[Hold (reason (Text 400))]
[Sell (reason (Text 400))])
(define analyst-ctx (filing)
(list (Message [role System] [content "You are an equity analyst."])
(Message [role User] [content filing])))
;; A script: what the model says, how big the reply is, how long it takes.
(script sell
[analyst (reply "{\"tag\":\"Sell\",\"reason\":\"Margins are compressing.\"}")
(out 20) (latency 2s)])
(under ([script sell] [cost $1.00] [time 1min])
(check-expect (ask claude Verdict (analyst-ctx "FY2025 10-K") 'analyst)
(Sell "Margins are compressing."))
;; the same call, measured: at most a cent and two seconds
(check-within (ask claude Verdict (analyst-ctx "FY2025 10-K") 'analyst)
([cost $0.01] [time 2s])))
2. Handling Failure
If a reply doesn’t fit the type, or the provider is down, or the money or time runs out,
ask produces a failure: one of Invalid,
ToolError, Refused, OverBudget, PastDeadline
or Raised. Agents fail often, so failure is part of the language rather than a crash.
(grant claude (model [in 3] [out 15] [ceiling 8000]))
(datatype Verdict
[Buy (reason (Text 400))]
[Hold (reason (Text 400))]
[Sell (reason (Text 400))])
(define ctx () (list (Message [role User] [content "FY2025 10-K"])))
(define analyst (model)
(catch (ask model Verdict (ctx) 'analyst)
err
(Hold "analysis unavailable")))
;; The model replies with prose where a Verdict was required.
(script garbled
[analyst (reply "I think you should sell.") (out 8) (latency 1s)])
(under ([script garbled] [cost $1.00] [time 1min])
(check-fail (ask claude Verdict (ctx) 'analyst) (Invalid _))
(check-expect (analyst claude) (Hold "analysis unavailable")))
3. The budget can never be overspent
Before calling the model, ask reserves the worst case: everything it
is about to send, plus the largest answer the type allows. If the reservation doesn’t fit,
the model is never called and nothing is charged. This is why types carry size bounds — and
why the same program with an unbounded answer type is refused.
(grant claude (model [in 3] [out 15] [ceiling 8000]))
(datatype Verdict [Sell (reason (Text 400))]) ; bounded: reserves about $0.007
(datatype Loose [LSell (reason Text)]) ; unbounded: reserves about $0.12
(define ctx () (list (Message [role User] [content "FY2025 10-K"])))
(script both
[tight (reply "{\"tag\":\"Sell\",\"reason\":\"Margins fell.\"}") (out 20) (latency 2s)]
[loose (reply "{\"tag\":\"LSell\",\"reason\":\"Margins fell.\"}") (out 20) (latency 2s)])
(under ([script both] [cost $0.05] [time 1min])
(check-expect (ask claude Verdict (ctx) 'tight) (Sell "Margins fell."))
;; The same question with an unbounded answer type will not fit $0.05.
(check-fail (ask claude Loose (ctx) 'loose) OverBudget)
;; And refusing costs nothing at all: no money, no time, no call.
(check-within (catch (ask claude Loose (ctx) 'loose) e 0)
([cost $0] [time 0s])))
4. Agent Loop
This is the CP-Agent design: ask for the next step, run any code it writes, feed the result back,
stop when it says it is done. Because the answer is a choice, the loop is an ordinary
case. The enclosing budget guarantees it stops.
(grant claude (model [in 3] [out 15] [ceiling 8000]))
(grant py (kernel)) ; a stateful tool, from the host
(datatype Step
[Exec (code (Text 4000))] ; "run this and show me the output"
[Done (code (Text 4000))]) ; "this is my final answer"
;; A tool error becomes an observation the model can read. Nothing else does.
(define observe (py code)
(catch (call py exec code) e
(case e [(ToolError m) m] [_ (fail e)])))
(define react (model py ctx)
(case (ask model Step ctx 'react)
[(Exec code)
(react model py
(append ctx (list (Message [role Assistant] [content code])
(Message [role Tool] [content (observe py code)]))))]
[(Done code) code]))
(define cp-agent (model py task)
(budget ([cost $0.50] [time 10min])
(react model py (list (Message [role User] [content task])))))
(script cp
[react (reply "{\"tag\":\"Exec\",\"code\":\"x = 6 * 7\"}") (out 12) (latency 3s)
(reply "{\"tag\":\"Done\",\"code\":\"print(42)\"}") (out 16) (latency 3s)]
[py/exec (result "") (latency 1s)])
(under ([script cp] [cost $1.00] [time 5min])
(check-expect (cp-agent claude py "Compute 6 * 7.") "print(42)")
;; two asks at 3s and one tool call at 1s
(check-within (cp-agent claude py "Compute 6 * 7.") ([cost $0.05] [time 7s])))
5. Concurrency
A workflow is a set of named steps, and the dependencies come from the names
each step mentions. Any step whose inputs are ready runs immediately. Below, the same
four model calls take 40 seconds as a dataflow graph and 60 seconds forced into
“run two, wait, run two” — at identical cost. Concurrency saves time, never money.
(grant claude (model [in 3] [out 15] [ceiling 8000]))
(define step-ctx (name)
(list (Message [role User] [content (string-append "Do step " name)])))
;; c needs a and b; d needs only a. Mentioning a name creates the edge.
(define n-graph (m)
(workflow ([a (ask m (Text 20) (step-ctx "a") 'a)]
[b (ask m (Text 20) (step-ctx "b") 'b)]
[c (begin a b (ask m (Text 20) (step-ctx "c") 'c))]
[d (begin a (ask m (Text 20) (step-ctx "d") 'd))])
(list a b c d)))
;; The same work, forced into two phases.
(define n-graph-par (m)
(let* ([ab (par (ask m (Text 20) (step-ctx "a") 'a)
(ask m (Text 20) (step-ctx "b") 'b))]
[cd (par (ask m (Text 20) (step-ctx "c") 'c)
(ask m (Text 20) (step-ctx "d") 'd))])
(list (. ab fst) (. ab snd) (. cd fst) (. cd snd))))
;; `elapsed` is six lines of µNorman, built on the (remaining) observer.
(define elapsed (thunk)
(let* ([t0 (. (remaining) time)] [_ (thunk)] [t1 (. (remaining) time)])
(- t0 t1)))
(script dag
[a (reply "\"A\"") (out 2) (latency 10s)]
[b (reply "\"B\"") (out 2) (latency 30s)]
[c (reply "\"C\"") (out 2) (latency 10s)]
[d (reply "\"D\"") (out 2) (latency 30s)])
(under ([script dag] [cost $1.00] [time 5min])
(check-expect (n-graph claude) (list "A" "B" "C" "D"))
(check-expect (elapsed (lambda () (n-graph claude))) 40s) ; critical path
(check-expect (elapsed (lambda () (n-graph-par claude))) 60s)) ; fork and join
6. Laws
Every construct has algebraic laws, so you know which rewrites preserve meaning. Just as useful
are the non-laws. A compiler would happily merge two identical
asks into one — and change the answer, because a model is a relation, not a
function. µNorman therefore never caches a model answer automatically.
(grant claude (model [in 3] [out 15] [ceiling 8000]))
(datatype Verdict
[Buy (reason (Text 400))]
[Sell (reason (Text 400))])
(define ask-v ()
(ask claude Verdict (list (Message [role User] [content "FY2025 10-K"])) 'analyst))
(script two-answers
[analyst (reply "{\"tag\":\"Buy\",\"reason\":\"One.\"}") (out 10) (latency 2s)
(reply "{\"tag\":\"Sell\",\"reason\":\"Two.\"}") (out 10) (latency 2s)])
(under ([script two-answers] [cost $1.00] [time 1min])
;; NOT a law: common-subexpression elimination. Two asks, two answers.
;; An optimizer that rewrote this to (Pair [fst a] [snd a]) would be wrong.
(check-expect (let* ([a (ask-v)] [b (ask-v)]) (Pair [fst a] [snd b]))
(Pair [fst (Buy "One.")] [snd (Sell "Two.")]))
;; A law that does hold: catching a failure only to re-raise it changes
;; nothing at all -- not the value, not the money, not the clock, not the trace.
(check-equiv (catch (ask-v) x (fail x)) (ask-v) 'exact))
What µNorman guarantees
- The budget is never overspent, even with many steps running at once. The interpreter checks this invariant in every state it reaches.
- A model can’t hand your program a tool. Authority comes only from the host.
- No answer is accepted after the deadline.
- Concurrency doesn’t change answers. A workflow gives the same result as running its steps one at a time.
- Every model call and tool call is recorded in a trace. That is what makes testing, replay and cost profiling possible.
Status
µNorman is a research prototype, designed in the style of Norman Ramsey’s Programming Languages: Build, Prove, and Compare and his Seven Lessons in Program Design. It is not ready to build a product on, and this section says why rather than leaving you to find out.
What works
The language, its semantics, its laws, and an interpreter that runs them. 127 unit tests, 13 laws checked on 150 random scripts each, and a suite of deliberately wrong tests that must fail.
It talks to a real model
ask works against the Anthropic API. The first real call confirmed the schema, the typed answer and the budget guarantee.
No real tools yet
call is scripted. There is no Python kernel or filesystem host, so an agent that uses a tool can’t yet run end to end against a real model.
No static checking yet
The guarantees above are enforced when a program runs. A type and effect system is designed but not built; it is what would move the checks to compile time.
A very small standard library
There is no number-to-string, no substring, no way to render a value back into a prompt. Real programs hit this within minutes.
Performance is a hypothesis
Concurrency, hard budgets and typed answers should make agents cheaper and more predictable. That has been shown in simulation, never measured against a real workload.
The answer’s type is sent with every ask as a JSON schema, which accounts for a significant portion of the input tokens. A type costs money twice: bounding an answer shrinks the reservation, but adding a constructor or a field makes every ask dearer on input, for as long as the program runs. That falsified one of the algebraic laws, which had priced an ask without its schema.