<!DOCTYPE html>
<html>
<head>
<title>Tarn Barford</title>
<meta charset="utf-8"/>
<link rel="icon" type="image/x-icon" href="/favicon.ico">
<link href="/style.css" media="screen" rel="stylesheet" type="text/css" />
<link rel="alternate" type="application/atom+xml" title="Journals of Tarn Barford" href="/atom" />
<link href="/highlight.css" media="screen" rel="stylesheet" type="text/css" />
<link href="/highlight-console.css" media="screen" rel="stylesheet" type="text/css" />
</head>
<body>
<div id="container">
<div id="header">
<div id="header">
<p>From the <a href="/journal">Journals</a> of <a href="/">Tarn Barford</a></p>
<h1>
vim-slime
</h1>
<p>
Mar 26, 2012
</p>
</div>
</div>
<div id="post_content">
<html><body><p>Today I found the awesomeness that is <a href="https://github.com/jpalardy/vim-slime">vim-slime</a>, it's been an exciting day for me.
<a href="http://common-lisp.net/project/slime/">Slime</a> is the "The Superior Lisp Interaction Mode for Emacs", I can almost hear the emacs crowd laughing.</p>
<p>For those that use vim and haven't used Slime, vim-slime or <a href="https://github.com/vim-scripts/VimClojure">something similar</a>, this is why it's awesome:</p>
<p><strong>Text can be sent from any process to the stdin of a <a href="http://www.gnu.org/software/screen/">gnu screen</a> or <a href="http://tmux.sourceforge.net/">tmux</a> session.
The process in this case is vim and the screen/tmux session is a terminal</strong>.</p>
<p>Screen is a <a href="/journal/oh-screen-where-have-you-been">really neat</a> terminal multiplexer (you can run multiple terminals in a terminal window).
The multiplexed shell processes are children of the screen process, which itself is not a child of the terminal window process.
This means a screen process and its child processes keep running if you close the terminal window.
Later you can re-connect to it, this is what makes vim-slime possible.</p>
<p>Here is an screen shot, on the left is me in gVim writing some awful Clojure <a href="#footnote-1">[1]</a>.
On the right is a screen buffer in which I started a Clojure REPL.
When I want to try run some code I can send any vim text selection to the REPL in a keystroke (or two).</p>
<p><img alt="vim slime screenshot" src="screenshot.jpg"/></p>
<p>It doesn't have to be a Clojure REPL either, we can send anything to a screen shell.
We could run git commands, find, grep, sed, etc. Like with the Clojure REPL we can even interact with any terminal programs that use STDIN.</p>
<p>This concept can be taken even further,
You can even connect to a tmux session over SSH and share a terminal or a <a href="http://remotepairprogramming.com/remote-pair-programming-with-tmux-and-vim-the">terminal program like vim to do remote pairing</a>!</p>
<p>Hopefully remote pairing is the topic of my next post as there are a couple geographically distant people I know who are keen to do some pair hacking.
I stand to learn a lot!</p>
<p><a name="footnote-1">[1]</a>
I learnt almost everything I know about Lisp from <a href="http://www.ccs.neu.edu/home/matthias/BTLS/">The Little Schemer</a>. Great book.</p></body></html>
</div>
<div id="comments">
</div>
</div>
<div id="footer">
<p> </p>
<p>Questions, comments, suggestions? Email me, <a href="mailto:tarn@tarnbarford.net">tarn@tarnbarford.net</a> (<a href="/pgp.txt">public key</a>)</p>
<p> </p>
</div>
</body>
</html>
<!DOCTYPE html>
<html>
<head>
<title>Tarn Barford</title>
<meta charset="utf-8"/>
<link rel="icon" type="image/x-icon" href="/favicon.ico">
<link href="/style.css" media="screen" rel="stylesheet" type="text/css" />
<link rel="alternate" type="application/atom+xml" title="Journals of Tarn Barford" href="/atom" />
<link href="/highlight.css" media="screen" rel="stylesheet" type="text/css" />
<link href="/highlight-console.css" media="screen" rel="stylesheet" type="text/css" />
<style>
#swipe-canvas {
position: relative;
width: 900px;
height: 300px;
}
#swipe-results {
font-size: 30px;
padding-left: 50px;
padding-left: 50px;
}
#swipe-results ul {
margin: 0px;
padding: 0px;
}
#swipe-results li {
float: left;
background-color: #DDDDDD;
list-style-type: none;
padding: 10px;
margin: 5px;
border-radius: 5px;
}
#swipe {
position: relative;
}
#swipe-loading {
position: absolute;
height: 50px;
width: 300px;
top: 85px;
left: 300px;
background-color: darkgray;
border-radius: 10px;
text-align: center;
padding-top: 20px;
border: black;
border-width: 5px;
}
</style>
</head>
<body>
<div id="container">
<div id="header">
<div id="header">
<p>From the <a href="/journal">Journals</a> of <a href="/">Tarn Barford</a></p>
<h1>
Swipe Keyboard
</h1>
<p>
Apr 06, 2014
</p>
</div>
</div>
<div id="post_content">
<html><body><p>When I first tried a <a href="http://www.swype.com/">Swype</a> keyboard I was impressed how effective it
was. Even though I don't use the feature on my phone I was interested in how it
could be built, so I <a href="https://github.com/tarnacious/swipe-keyboard">implemented this otherwise useless swipe-able keyboard</a> below. It probably doesn't work on mobile devices, but works on modern
browsers with mouse pointers (although I've only really tried Chrome and
Firefox).</p>
<div id="swipe">
<canvas height="300px" id="swipe-canvas" width="900px"></canvas>
<div id="swipe-results"></div>
<div style="clear: both"></div>
<h2 id="swipe-loading">Loading<noscript>Javascript is Required</noscript></h2>
</div>
<p>I initially tried to solve this using the technique Peter Norvig famously uses
in his <a href="http://norvig.com/spell-correct.html]">spell checker</a>. He takes a sequence of characters and
generates a set of word candidates by adding, removing and swapping characters
in the original sequence, the generated candidates are removed if they are not
found a dictionary. This can work but to be effective too many combinations
need to be generated.</p>
<p>If the dictionary is indexed into a <a href="http://en.wikipedia.org/wiki/Trie">trie</a> the number of combinations
generated can be reduced significantly by traversing the trie and only
generating valid letter combinations. This is a pretty bare implementation of
that, it requires: </p>
<ul>
<li>The first and last characters of the initial sequence are used </li>
<li>Intermediate characters in the initial sequence can be repeated or discarded </li>
<li>No characters are added or swapped</li>
</ul>
<p>Basically, if you swipe through all the characters in a word in order, then the
word will be found if it is in the index regardless how many characters are
swiped in between. It is surprisingly quick and effective.</p>
<p>This implementation uses <a href="https://raw.github.com/first20hours/google-10000-english/master/google-10000-english.txt">these 10000 words</a>, I intended to use digital
books but never got around to it as these words demonstrate the concept well
enough.</p>
<p>This is the first thing I've written in <a href="https://github.com/clojure/clojurescript">ClojureScript</a> or
<a href="https://github.com/clojure/clojurescript">Clojure</a>, so my code my vary from non-idiomatic to shamblolic. I
initially used a <a href="http://clojuredocs.org/clojure_core/clojure.zip/zipper">zipper</a> to build the trie with immutable data
structures, but found the indexing took to long with my zipper implementation
so I <a href="https://github.com/tarnacious/swipe-keyboard/commit/6edd7b26e78121fbe8586b3f0ef54ca8277d9e32">switched to using native Javascript maps</a>.</p>
<p>I found that <a href="https://github.com/clojure/core.async">core.async</a> library is really awesome, the <a href="http://docs.closure-library.googlecode.com/git/index.html">Google
closure library</a> and <a href="https://developers.google.com/closure/compiler/">compiler</a> integration with <a href="http://leiningen.org/">Leiningen</a> the <a href="https://github.com/emezeske/lein-cljsbuild">cljsbuild plug-in</a> to be impressive. My main pains
were the slow JVM start-up time, the advanced closure compiler build of the web
worker script fails silently when run (but the main script works fine when
compiled with the advanced compiler), and at times I felt some compile time
type checking would be nice.</p>
<p>I would like to extend this experiment to index the word occurrence counts and
proceeding word counts in original text and rank the found words as most
likely. Support casing, umlauts, special characters, spelling correction and
compound words in the indexing and lookup. I think a live lookup while swiping
would also be possible.</p>
<p>Overall this was fun, turned out OK I think, and was a great learning
experience.</p></body></html>
</div>
<div id="comments">
</div>
</div>
<div id="footer">
<p> </p>
<p>Questions, comments, suggestions? Email me, <a href="mailto:tarn@tarnbarford.net">tarn@tarnbarford.net</a> (<a href="/pgp.txt">public key</a>)</p>
<p> </p>
</div>
<script src="swipe.js" type="text/javascript"></script>
</body>
</html>
An agent that books invoices, answers questions over a database or fills in spreadsheets is more than a language model. It is a workflow: the model, the instructions it is given, the tools it may use and the steps it takes. Every part of it can be changed, and every change has a price. Measuring a workflow on a fixed set of cases is the statistical counterpart of a regression test, and it is what makes a change safe to deploy. This article shows how Dvergr, our open-source system for running and measuring agents, measures such workflows, what it found when we measured ten models on two public benchmarks, and why the most useful thing a measurement finds is often not the best model but a better workflow.
The question behind a leaderboard
Suppose a firm wants an agent to answer its staff’s questions about the company database. It has to choose a model, write the instructions, and decide what the agent may do: write one query and stop, or look at the data first, try a query, read the error and try again. A public leaderboard answers a narrower question, which model scored highest on someone else’s cases, with someone else’s instructions, on one run.
Dvergr answers the firm’s question directly. It takes a set of cases, each a task with the answer a person recorded for it, and a set of candidates, each a model together with its instructions and tools. It runs every candidate on every case and reports, for each candidate:
how often it is right, with a range that says how much that number could move on another set of cases of the same kind;
where it fails, check by check;
what a correct answer costs at the model’s list price;
how long it takes.
A firm that deploys an agent needs what software teams get from regression tests: a fixed set of cases with recorded answers, run again on every change. An agent’s output varies from run to run, and a provider can change the model behind a name it keeps, so one passing run says little. The result of a run is therefore a rate with an interval rather than pass or fail, and a change is judged by comparing it with the previous version on the cases where the two disagree.
We tried this on two public benchmarks, where anyone can check the cases and the answers. SpreadsheetBench asks for changes to Excel workbooks. BIRD asks questions in plain language about real databases, to be answered with a query.
The models come in families, named here from cheapest to most expensive: OpenAI’s Luna, Sol and Astra, in generations 5.6 and 6 (6.1 for Sol); Anthropic’s Claude Haiku, Sonnet and Opus 5.5; and two models with published weights, GLM and DeepSeek.
Results at a glance
Correctness with its 95 % interval against cost per correct answer, at list price, measured 2026-10-01 to 2026-10-10. BIRD: seven models writing SQL for SQLite on the same 100 questions. SpreadsheetBench: eight models on the same 59 tasks. Overlapping intervals are ties. On BIRD, GPT-6.1 Sol and GPT-6 Astra score highest, the first models we measured to rise above the rest, though their intervals still overlap the others'; on SpreadsheetBench the models tie, at costs that differ by more than a hundredfold.
Each dot is one model. Its bar shows how far the score could move if we had picked a different 100 questions of the same kind (59 tasks on SpreadsheetBench), the way a poll of 100 people would come out differently with another 100. Each model answered each question once, and the bar follows from that count alone, which is why it is nearly the same length for every model: with 100 questions and scores around two thirds, a score can move by about nine points either way. Where two bars overlap, the measurement cannot tell the two models apart, and the choice between them comes down to cost and time, on the horizontal axis.
Because every model answered the same questions, two models can also be compared more sharply. Most questions do not help, because both models got them right or both got them wrong. What decides is the questions where they differ. If model A was right and B wrong on twelve questions, and the reverse happened on only two, A is better even though their bars overlap; six against five says nothing. We give these counts with a p-value, the chance that two equally good models would split their differences at least this unevenly, and below 0.05 we treat a difference as real. Dvergr’s report gives the bars; the counts and p-values here we computed from the per-question results it records.
The tables also give the median time per case. Time depends on how busy the provider was that day, so it compares candidates run together better than candidates run on different days.
How a measurement runs
A measurement. Certified cases enter an experiment room; each attempt runs in its own fork and is graded there; only the verdict is kept.
A benchmark enters Dvergr as a case pack: the cases, their recorded answers, and a checker that decides whether an answer is right. Before any model runs, the pack is certified. Certification lists every case that cannot grade an answer and says why: an id that is missing or appears twice, an empty answer, the same inputs recorded with different answers, or a recorded answer the checker itself does not accept. Those cases are set aside. On SpreadsheetBench, certification kept 390 of the 400 tasks without spending any model tokens.
The measurement itself takes place in a room, Dvergr’s unit of shared state: its files, its database and its history. For each candidate and case, Dvergr makes a fork of the room, a private copy that costs almost nothing to make because it shares everything until something changes. The candidate works in its fork: the files and the workbook it changes belong to that copy, and a database it may only read is reached for each query through a fresh read-only connection or an unchangeable snapshot, so nothing it does reaches the original data or the candidates working beside it. The checker grades the attempt’s result (the files in the fork, the query it submitted, or its edits replayed on a fresh copy of the workbook), the verdict is recorded in the room, and the fork is thrown away. Dvergr’s own share of each case, the fork included, is a fraction of a second, against the ten to seventy seconds a model takes to answer.
The statistics depend on these forks. Because every candidate starts from the same state, two candidates’ answers to a case can be compared directly, which is what the paired comparison needs. Because no attempt can change what another sees, a candidate is graded on its own work. And because a fork shares everything that has not changed, running every candidate on every case stays cheap. This follows from how Dvergr keeps state: the room’s files and its database are forked with copy-on-write branches across git and Datahike, so a fork records only what it changes and never writes to the original.
Four rules keep the numbers honest:
The same cases for everyone. Every candidate answers every case, which is what makes the paired comparison possible.
A broken connection is not a wrong answer. When the provider is down or a request times out, the case is run again. A model that answers wrongly, or gives no answer, gets its verdict.
List price. Models are billed per token, a piece of a word, read or written. Cost is what the same tokens would cost through the provider’s public price list, whether a subscription or the firm’s own key paid for them, with text the provider has already seen recently (cached input) at its lower rate.
A changed setup is a new experiment. The cases, candidates, instructions, checker and software versions together define an experiment. An interrupted run resumes where it stopped, but if any of them has changed, running it again starts a new experiment instead of mixing two.
Does the agent need to explore?
The candidates above are models inside Dvergr’s own workflows. On BIRD, the model may run queries in SQL, the standard database query language, see their results or errors, and submit an answer when it is satisfied, within twenty turns. On spreadsheets it may read cells, write values and formulas, see what they compute, and submit.
A firm should know how much of a score is the model and how much is the workflow around it. An agent that explores uses more tokens than one that answers at once, and it is worth knowing what those tokens buy. So we also ran each benchmark’s own reference setup, with the same model on the same cases. These runs are not Dvergr workflows: they use each benchmark’s published code to build the prompts and grade the answers, and Dvergr only connects them to the model, so that both setups reach it the same way. The scripts, with the steps to reproduce each number below, are in the repository under benchmarks/reference/.
On BIRD, one instruction did most of the work
BIRD’s reference setup is a single prompt: the list of the database’s tables and columns (its schema), the question and a hint, and a request to write the SQL query. The model answers once, without seeing any data. We ran it as published, then with one line added, return exactly the columns the question asks for. Graded with BIRD’s own evaluation script, on the same 100 questions:
Workflow
GPT-5.6 Luna
Claude Haiku 5.5
Tokens per question (Luna)
BIRD’s prompt, one reply
51
53
1,200
the same, plus the line about columns
61
64
1,200
Dvergr: run queries, see results and errors, then submit
65
69
8,200
Two questions show what each part contributes. Question 781 asks for the heights of the heroes whose eye colours are amber. With BIRD’s prompt, Luna returned each hero’s name beside the height. The heights were right, but BIRD compares the returned table as a whole, and a table with an extra column is wrong. With the line about columns it returned the heights alone and passed. Of the fifteen questions Luna answered on our workflow and missed with BIRD’s prompt, seven failed for no other reason than such an extra column, and the one line is worth ten points for both models, a difference too large for chance (p = 0.013 and 0.003).
Question 817 asks for the race of the blue-haired male superhero, and its hint says the colour is written 'blue' and the gender 'male'. In the database they are written 'Blue' and 'Male', and SQLite compares text exactly. The one-reply query followed the hint and found nothing. In Dvergr’s workflow Luna wrote a query that ignores capitals, ran it, saw three heroes, and submitted the query for their races. That is the kind of question where looking at the data pays.
Exploring the data adds another four or five points. With 100 questions that could be chance (p = 0.34 and 0.30), and it costs seven times the tokens. For these questions the cheapest good workflow is a single reply with the right instruction, at about a third of the cost per correct answer. The agent that explores scores higher, but 100 questions cannot show that the gain is real, and it costs more.
We had written that line into our own workflow from the start, because we knew how BIRD grades. Until we measured it, we did not know it was most of what our workflow added.
On spreadsheets, the recalculating program decided the score
On SpreadsheetBench the scores depended less on the model than on how the answers were checked.
The benchmark’s reference setup shows the model the instruction and the first five rows of each sheet, and asks for a short Python program that edits the workbook. Its multi-round version runs the program and shows the model the output or error, for up to five turns. We ran both versions with GPT-5.6 Luna on our 59 tasks, each program in a Docker container that can see only its own input file, and graded them with the benchmark’s own comparison code. Dvergr’s workflow writes no Python: its agent edits the workbook directly, through a spreadsheet engine in Dvergr’s sandbox.
SpreadsheetBench’s grader reads each answer cell’s stored value, the result a spreadsheet program saved with the formula, and never computes a formula itself. A file written from Python stores formulas without values, so the benchmark’s instructions add a step before grading: open every workbook in LibreOffice (on Linux and macOS) or Excel (on Windows) and let it recalculate. We followed the instructions, using LibreOffice on Linux.
One task shows what happened. Task 56563 has a column of monthly amounts, one of which is an error value, and asks for a formula that adds them up anyway. Both workflows wrote the same correct formula into the answer cell, =AGGREGATE(9,6,C2:C13), which sums the range and skips cells holding errors. Its value is 77,772.6. We graded that one cell two ways:
Graded by
What it found in the cell
Verdict
the grader after LibreOffice recalculated, as the benchmark’s instructions say
#NAME?: LibreOffice recognises functions added to Excel since 2010, such as AGGREGATE, only when the file marks them with a prefix (_xlfn.AGGREGATE), and neither workflow wrote it
wrong
Dvergr’s spreadsheet engine
77,772.6, as does LibreOffice once the prefix is written
right
The same answer was wrong or right, depending on which program computed it. Across the 59 tasks:
Workflow
Recalculated by LibreOffice (the benchmark’s instructions)
Tokens per task (median)
Median time
SpreadsheetBench, one reply with code
34
2,200
15 s
SpreadsheetBench, up to five turns with execution
41
4,300
17 s
Dvergr: read and write cells, see what they compute, submit
41 as written, 48 as Excel stores them
5,700
24 s
Graded the benchmark’s way, after LibreOffice’s recalculation, the multi-round workflow scores 41, and Dvergr’s workbooks as written also score 41; each is right where the other is wrong on eight tasks, so on that footing they tie. Written with Excel’s prefix, Dvergr’s workbooks score 48. Graded by Rechentafel, the open-source spreadsheet engine Dvergr’s agent works in, they score 51. Grading with our own engine would prove little if the engine had not been checked against Excel. The benchmark’s recorded answers are the values Excel saved, and for 390 of the 400 tasks Rechentafel recalculates each recorded answer to exactly that value, which is how certification works on this benchmark. The workflows differ less than the graders do. A firm measuring a spreadsheet workflow should grade with the program its people use, or with one checked against it.
Improving a workflow by measuring it
The line about columns is the kind of change a measurement should find, not one an author should have to know in advance. Dvergr is built so that variants of a workflow can be written and measured like any other candidate.
A workflow in Dvergr lives in its room as ordinary files: the task and its instructions, the checker, the cases, and any code the agent runs in its sandbox. Because they are files in the room, an agent working there can read them and write a variant: a sentence added to the instructions, a look at the data before the answer, a different model. Each variant is a candidate like any other, measured in its own forks on a set of tuning cases. A variant that does better is then measured on cases it has not seen, so that one tuned to the quirks of one set of questions does not pass for an improvement. A variant’s checker is trusted only after the host’s administrator promotes it, through an administrator’s connection. The promotion is recorded outside the room’s files, where no agent can write, so an agent cannot raise the trust of its own checker.
We took one step of this loop by hand on our Datalog workflow for BIRD. Datalog is the query language of Datahike, the database Dvergr keeps its own state in, and the workflow answers BIRD’s questions with it instead of SQL. After reading the failures of one run we added two sentences to its description of the language, and measured the change on another 100 questions that run had not seen:
Workflow
Correct (of 100)
Input tokens per question
SQL on SQLite
59
8,880
Datalog, revised description
63
14,042
Datalog, previous description
58
16,684
The revision answers five more questions with less input, but at 100 questions that could be chance (p = 0.23). That is the honest result of one step of tuning: a promising change, to be confirmed on more cases before it is adopted. Today a person or an agent proposes each variant and Dvergr measures it. Letting an agent in the room run the whole loop, propose, measure and hand the winner to the owner, is what we are building next. The rule for adopting a variant is the one a team applies to a code change: merge it when the cases show it is not worse.
SpreadsheetBench in detail
This section and the next give the full results behind the chart, model by model. These numbers use our workflow and our spreadsheet engine, on 59 of the 390 certified tasks, so they are not comparable to the public leaderboard, where agents work with other tools.
Model
Tasks
Passed
95 % interval
Cost per pass
Median time
GPT-5.6 Luna
59
51
76–93 %
$0.0026
24 s
GPT-5.6 Sol
59
52
78–95 %
$0.033
27 s
Claude Opus 5.5 (Claude Code subscription)
59
52
78–95 %
$0.062
17 s
Claude Haiku 5.5 (Claude Code subscription)
59
47
68–88 %
$0.0046
22 s
Claude Sonnet 5.5 (Claude Code subscription)
59
46
66–87 %
$0.33
19 s
GPT-6 Luna
59
48
70–90 %
$0.0008
10 s
GPT-6.1 Sol
59
55
85–98 %
$0.010
9 s
GPT-6 Astra
59
55
85–98 %
$0.050
9 s
Paired with GPT-5.6 Luna, none of the GPT-5.6 and Claude models is measurably different. Opus is right where Luna is wrong six times and the reverse five times; Sol three and two; Haiku four and eight; Sonnet three and nine (the closest, p = 0.15). What differs is the cost. Per passing task Haiku costs about twice as much as Luna, Sol thirteen times, Opus twenty-four times and Sonnet over a hundred times, because Sonnet sometimes wrote very long answers: its slowest tenth of tasks took over nine minutes.
The GPT-6 models, run on the same tasks a few days later, are the first to move the top. GPT-6.1 Sol and GPT-6 Astra pass 55 each and disagree on only four tasks, two each way, so Astra costs five times as much as Sol for the same answers. Sol 6.1 is measurably better than GPT-6 Luna (eight tasks to one, p = 0.039) but not than GPT-5.6 Luna (five to one, p = 0.22). They were also about three times faster, partly because of the models and partly because of the day.
Every task but one was solved by at least one model, so most failures belong to a model, not to the task.
BIRD in detail
BIRD’s questions cover 11 databases. We can answer each question three ways: SQL on SQLite, the database BIRD ships with; SQL on Datahike, through its PostgreSQL-compatible layer pg-datahike; and Datalog on Datahike. Even before any model runs, 1372 of BIRD’s 1534 reference queries return exactly SQLite’s rows on pg-datahike.
With GPT-5.6 Luna on 100 fresh questions held out from tuning:
Engine
Correct (of 100)
95 % interval
Cost per correct answer
Input tokens per question
Median / slowest tenth
SQL on SQLite
65
55–74 %
$0.0023
7,839
10 s / 23 s
SQL on Datahike (pg-datahike)
65
55–74 %
$0.0026
10,136
10 s / 41 s
Datalog on Datahike
64
54–73 %
$0.0044
15,061
13 s / 42 s
The three engines answer equally well: each pair disagrees on seven to ten questions, split almost evenly. They differ in cost. In Datalog the model looks at the schema more before it answers, so a correct answer costs about twice as much as in SQL. Which query language to give an agent is a cost decision, and the measurement says how large it is.
Seven models on the same 100 questions:
Model
SQL on SQLite (of 100)
Datalog (of 100)
Cost per correct answer, SQL / Datalog
Median time, SQL
GPT-5.6 Luna
65
64
$0.0023 / $0.0044
10 s
Claude Haiku 5.5
69
68
$0.0026 / $0.0047
9 s
DeepSeek V4.1 Flash (open weights)
66
64
$0.0074 / $0.0195
8 s
GLM 5.3 Flash (open weights)
64
58
$0.0042 / $0.0106
6 s
GPT-6 Luna
68
60
$0.0009 / $0.0014
5 s
GPT-6.1 Sol
74
71
$0.015 / $0.019
8 s
GPT-6 Astra
75
70
$0.067 / $0.091
6 s
Among the first four models, no two are measurably different. Any two disagree on 10 to 18 questions; the most lopsided pair (GLM’s Datalog against Haiku’s, p = 0.013) is no more lopsided than chance alone produces somewhere among 28 comparisons.
GPT-6.1 Sol and GPT-6 Astra answer 74 and 75 in SQL, against 65 for GPT-5.6 Luna (eleven questions to two and twelve to two), and tie with each other. Counting all the comparisons we made, even that is not quite conclusive, but it is the first gain we have seen, and it comes within about five questions of the 80 that BIRD’s recorded answers allow (see below). Again Astra costs about five times as much as Sol per correct answer.
The two open-weight models (models anyone may download and run on their own machines), served by Fireworks, answer as many questions as GPT-5.6 Luna and Claude Haiku, at two to three times the cost per correct answer, because they used more tokens and none of their input was served from a cache. For a firm whose data may not leave its own machines, that is the useful result: a model it can run itself answers these questions as well, and its cost is then the firm’s hardware, not a list price.
What the benchmarks get wrong
A benchmark is only as good as its recorded answers and its grader, and both benchmarks have faults that change scores.
On BIRD, twenty of the 100 questions were answered correctly by none of the first four models in either language. On most of them the models agree with each other and not with the recorded answer:
Question (BIRD id)
Recorded answer
What the models answered
atoms of molecule TR346 and its bond types (309)
a query for molecule TR000
molecule TR346
the German type of a card (482)
the English type
the German type
patients with abnormal CRP and no recorded data (1256)
208, the number of lab rows
25 patients; the only answer graded right, GPT-6 Luna’s in Datalog, repeated the mistake
expenses of the budget with the lowest remaining (1365)
one row, cut off by LIMIT 1
both expenses of that budget
difference of two percentages (1458)
12.12
0.1212, following the hint’s formula, which has no ×100
eye colours of Marvel heroes by popularity (728)
colour, count and a rank column
colour and count
We count about ten questions whose recorded answer is wrong, four more whose hint contradicts it, and the rest a matter of which columns or which wording. On these 100 questions the best attainable score is about 80, not 100. Paired comparisons survive this, since every candidate loses the same questions, but the differences shrink, and a model that answers the question as asked gets no credit for it. Our grader agrees with BIRD’s own on all 200 SQL answers we checked, so the scores above are the benchmark’s own.
On SpreadsheetBench:
The published evaluation script cannot score the published tasks. On the repository’s main branch it compares each task’s input workbook with the recorded answer (the line that reads the model’s output is commented out), and its file names are those of an earlier release. Every published score on the 400 verified tasks comes from a modified script that is not part of the benchmark.
Formulas are graded as empty unless the workbook is recalculated first, as described above. The benchmark’s instructions include that step, but the program that recalculates changes the score, and the step also recalculates the recorded answers, so after it they are that program’s results too.
One task can never pass. The answer position of task 45944 contains spaces, and the checker fails on it whatever the answer.
Some recorded answers are wrong or misfiled. Task 118-50’s answer misses a pair its own rule finds in the input, and the two strongest models were graded wrong for finding it. Task 42930’s answer file carries another task’s number.
Certification sets aside ten of the 400 tasks for reasons like these before any model runs.
Checking our own measurements
The setup that measures is code too, and it needs the same tests as the workflows it measures. Before publishing, four AI review agents read our benchmark code, documentation and stored results, each with one question: whether every benchmark runs the same way, whether the forks isolate each attempt, whether the documentation matches the code, and whether the numbers are right. They found nine problems, among them a provider outage counted as the model’s failure and cached input billed at its full price. Each fix came with a test that fails without it, and the affected runs were repeated or withdrawn.
What these numbers cannot tell you
Each model answered each case once, so a few answers could change on another run; the intervals say how much that matters. The numbers describe these cases with these workflows, and a different prompt, tool or model version can change them, which is why an experiment records all three. A tie at 59 tasks says what 59 tasks can separate, not that two models are equal. As with tests, coverage sets what can be found: with 100 cases, a regression of a few points can go unnoticed. And a public benchmark is not your workflow: it shows that the method works where others can check it, not how a model will do on your cases.
Where the method comes from
Seen this way, a workflow is a probabilistic program: a program whose output is drawn from a distribution, and measuring it is inference about how often that output is right. The Jeffreys interval in our reports is a Bayesian estimate of that rate. The view comes from probabilistic programming, which Christian Weilbach worked on in his doctoral research in Frank Wood’s group, with the Anglican language and the Daphne compiler (TMLR 2025). It continues in Foerster, our library for inference over forkable worlds, where alternatives are forks of the same state, as candidates are here. Dvergr’s measurements do not use Foerster yet. Two steps it would allow are stopping a comparison as soon as the evidence decides it and spending more cases where two candidates are close.
Try it on your own cases
The benchmark that matters for a firm is its own history: past bookings, coded invoices or resolved tickets, each with the outcome a person recorded. Dvergr is open source and runs on your own machine:
Install and connect. Clone the repository (it needs the Clojure CLI and babashka), and add bin/dvergr-mcp --profile bench as an MCP server in Claude Code or Codex. The models it compares can be used through a Claude Code or Codex subscription or any API key; the README lists the providers, the setup and where Dvergr keeps its state.
Turn your history into a case pack. Export a table of past cases as CSV, as Excel or DATEV write it, and name its columns: the id, the inputs, and the outcome with a rule for checking it (exact, a number within a tolerance, a date, a set). Certification reports which cases can grade an answer and why the others cannot, which is a first result in itself.
Run the comparison. Ask your assistant to run catalog_benchmark with the models you want compared, or, from a shell, turn the CSV into a pack with clojure -M -m dvergr.catalog.casepack-cli and run clojure -M -m dvergr.catalog.room-run <pack> --models … --cases 30. The report gives each model’s success with its interval, where it fails, the cost per correct result and the time.
Running benchmarks and case packs in the documentation give the details. A following article applies this to a year of bank bookings exported from DATEV.
If you would rather we built the benchmark with you, we run fixed-fee pilots: six weeks on one workflow of yours, ending with the benchmark and a report on which model, instructions and steps to run, at what cost. Send us a note.
Method
Models
GPT-5.6 Luna and Sol, GPT-6 Luna, GPT-6.1 Sol and GPT-6 Astra (Codex subscription); Claude Haiku 5.5, Sonnet 5.5 and Opus 5.5 (Claude Code subscription, each pinned to its exact version); GLM 5.3 Flash and DeepSeek V4.1 Flash (Fireworks). Each answers every case once.
Cases
BIRD dev, the held-out third of each database’s questions, fresh samples of 100; SpreadsheetBench verified, held-out tasks among the 390 certified
Grading
BIRD: the returned rows equal the recorded query’s rows; SpreadsheetBench: the answer cells equal the recorded workbook’s values
Statistics
Jeffreys 95 % intervals for success rates, as Dvergr’s report gives them; exact McNemar test on the cases two candidates disagree on, computed from the recorded per-case verdicts
Cost
list price per million tokens on the day of measurement, cached input at the cache rate
Reference setups
BIRD’s prompt from its published gpt_request.py and its evaluation script; SpreadsheetBench’s single- and multi-round drivers and its comparison code, the generated programs run in a Docker container; both reach the model through Dvergr’s model connection: benchmarks/reference/
Code and data
BIRD and SpreadsheetBench are public; Dvergr and both benchmark adapters are open source: github.com/replikativ/dvergr
For more than a decade the Clojure REPL most people saw in a terminal was
REPLy, the one behind lein repl. It
served us well, but it’s showing its age, and I kept wishing for a terminal
REPL that felt like CIDER, without the Emacs part. So I finally built one.
neorepl 0.1 is out today!
Why another REPL?
There’s no shortage of Clojure REPLs, so it’s fair to ask. Here’s what I was
after:
nREPL all the way down. neorepl only speaks nREPL, so it works with anything
that runs an nREPL server: Clojure, babashka, shadow-cljs, Basilisp,
ClojureCLR and so on. The smarts live on the server. When cider-nrepl is
there neorepl uses it, the same way CIDER does, and when it isn’t it falls
back to nREPL’s built-in ops, and then to evaluating a bit of code.
A REPL that’s pleasant to use: a proper line editor, highlighting,
completion, arglists as you type, pretty printing, history per project and
a Ctrl-C that actually interrupts the evaluation.
A CLI that’s just as good for scripts and coding agents. Every lookup the
REPL can do is also a command, with structured output and exit codes that
mean something.
Something small and modern. Clojure 1.10+, nREPL and JLine 4 are the only
dependencies, and the same code runs on the JVM and on babashka. The whole
thing is about 3,300 lines right now, and I’d like to keep it that way.
rebel-readline and bb repl --connect are both great, and I borrowed from
them freely, but neither tries to be an nREPL-first tool with CIDER’s tooling,
jack-in and a CLI for agents. That combination is the whole point of neorepl.
For humans
Here’s the REPL in action:
A few things worth pointing out:
Enter only submits the input once its forms are complete, so pasting a whole
function just works.
TAB completes, with the candidates’ types next to them, and the arglist of
the function you’re calling shows up after the cursor.
Values get pretty printed on the server and highlighted on the client.
Commands start with a comma, the CIDER way: ,doc, ,source, ,apropos,
,macroexpand, ,pst, ,reload, ,edit (which opens a definition in your
$EDITOR) and ,help for the rest. The Clojure reader sees a comma as
whitespace, so they can’t clash with real code.
Ctrl-C interrupts what’s running, and code that reads *in* gets its input,
something REPLy never quite got right.
Without a server to connect to, neorepl starts one for the project, the way
CIDER’s jack-in does, with cider-nrepl on the classpath, and stops it when you
exit. It knows Leiningen, the Clojure CLI, babashka and shadow-cljs projects.
ClojureScript works too. Start a ClojureScript REPL in the session (through
shadow-cljs or piggieback, which jack-in adds for you) and neorepl follows
along: evaluation, completion, docs and the rest switch to ClojureScript, and
:cljs/quit takes you back.
For agents
Coding agents are surprisingly good at Clojure, as long as they have a REPL
to talk to. Most of them drive CLIs much better than anything else, so neorepl
has a CLI that’s meant to be driven:
neorepl eval evaluates code on the project’s server. Calls share a
session, so *ns*, *1 and *e carry over from one call to the next, and
--session keeps separate lines of work apart.
--format json (or edn) gives you the values, the output and the
namespace as data.
--timeout interrupts the evaluation before giving up, so an infinite loop
doesn’t take the server down with it, and --max-output keeps a stray
(range) from eating the agent’s context.
neorepl check finds unbalanced parens, brackets and braces without a
server. That’s the most common way for an agent (or a human) to break a
Clojure file.
neorepl doc, source, apropos, complete, where, macroexpand and
reload do what the REPL’s commands do.
neorepl server starts a server that stays up, for editors, agents and
neorepl eval to share.
The exit codes are boring on purpose: 0 when all went well, 1 for an
evaluation error, 3 when there’s no server to talk to and 124 for a timeout
(like timeout(1)).
Playing nicely with nREPL
A new client is also a new way to break servers, or to be broken by them, and
the nREPL world has quite a few servers these days. Fortunately I’ve been
working on proof, a compatibility suite for
nREPL, which made neorepl a good test subject.
proof checks both sides of the conversation. proof proxy sits between a
client and a server and grades everything the client sends: every request has
an id, need-input gets answered with stdin in the right session, and so on.
proof serve is a small nREPL server that can behave like other servers in
very specific ways - sending only the last value of the code like Basilisp and
jank, dropping stderr, having no interrupt op, splitting output into one
message per character, and so on. I ran neorepl’s REPL and CLI through both,
against nREPL, babashka, Basilisp and jank, and through every quirk proof
serve knows about.
neorepl came out of it in decent shape, but not unscathed:
Code read from stdin (neorepl eval -) that then read *in* itself crashed
neorepl and left the evaluation waiting on the server for input that was
never coming.
neorepl’s errors could leave the shell’s prompt dangling at the end of a
line with servers that don’t end their error reports with a newline (hello,
jank!).
neorepl eval exited with 1 when it lost the server, which is the status
for code that failed.
All of those are fixed in 0.1. It also turned up a couple of things that are
nREPL’s fault rather than neorepl’s: need-input and the end of an
interrupted evaluation can go to the wrong connection when a session is used
from more than one. Those will be fixed in nREPL itself.
It works the other way around too. proof now checks servers with the exact
requests neorepl sends while starting its REPL and evaluating code, next to
the ones CIDER, Calva, Conjure, vim-fireplace and REPLy send, so a server that
breaks neorepl finds out from proof, not from a bug report. nREPL, babashka,
Basilisp and jank all pass those checks today.
Making the neighbours better
Building neorepl turned up a few problems elsewhere, which is one of my
favourite side effects of writing a new client. piggieback used to evaluate
only the first form you sent it and silently drop the rest, and it loaded the
ClojureScript compiler at startup even when nobody asked for a ClojureScript
REPL, which added almost half a second to the startup of every server with
ClojureScript on the classpath. The fixes for both
(piggieback#162 and
piggieback#161) will be in the
next piggieback release. The Node.js
REPL in ClojureScript itself doesn’t cope well with interrupted evaluations,
and that one is next on my list.
What’s next
0.1 is the foundation. Here’s what I’d like to tackle next:
a test runner with proper reports
a full-screen inspector and stacktrace browser
an MCP server and a ready-made skill for coding agents
It’s early days, so I’m sure you’ll find rough edges. Please file issues,
and tell me what you’d like to see next. And if you’re one of the people who
kept lein repl and REPLy going over the years - thank you! neorepl wouldn’t
exist without you.
I started building a useful personal agent. What makes it useful is that it runs on my computer, but this immediately raises concerns about data exfiltration. I use Apple reminders (among other things) to organise my life so I had it write a little script to be able to interact with those, which works great. One of the main things I want help with is getting a handle on my chaotic piles of todos. They are scattered all over the place and all mixed up. Part of the problem with my current system is that aspirational things are mixed up with hard deadlines, so on busy or unpredictable weeks things that aren’t essential just pile up and stay there and then the “today” list just keeps rolling everything over until there are dozens of things on it, which I couldn’t possibly get done in a single day if I tried. And then I’m just mentally keeping tabs on what is actually important or on a real deadline, which defeats the purpose of the system.
The problem is that since I had a baby every day is unpredictable. I used to have a pretty good handle on my life, but none of my old systems really work anymore. There’s no way to predict how my nights will go anymore so my energy levels are very inconsistent, and I can’t even know what the days will be like. Anyway sounds like I need a follow-up post on the woes of working parenthood. Point being, I need a new system and so far am having fun building an LLM-based one.
Like I said what makes the agent useful is that it runs on my computer. It has access to my data and the internet, and when it was just helping me with reminders it only had access to those, which I write, so I trust them.
The problem is that in the process of having it help me organize the reminders, I realized that the answers to most of the questions it was asking me could be found in my emails. It was still helpful and less overwhelming having a bot help me organize things, but it still needed a lot of information and context from me that it could have found in my emails if it’d had access. But the problem with giving a bot access to email is that emails come from other people. A way to inject a prompt to my bot is the missing piece of the lethal trifecta, so I need to find a way to do it safely. Turns out it’s just a hard problem, but there is some interesting research out there and I’m working on coming up with something that will work to safely allow the bot get answers from my email without a way to pass malicious or sensitive information it finds there forward to an agent that has ways to access the internet.
I’ve been stringing you along for years. Time for a (breaking?) change?
Shall I compare thee to a … string?
When porting Clojure for the JVM to the CLR, the question often arises of when a particular aspect of computation is intrinsic to Clojure or just an exposure of a feature of the JVM. For the former, I try to duplicate behavior; for the latter, I try to expose the corresponding CLR behavior. I have run into this not infrequently, particularly in the early days of the porting effort.
One area where it came up very early was with string comparison. And, frankly, I did not give it a lot of thought. Where ClojureJVM used java.lang.String.compareTo, I used System.String.CompareTo. In other words, I made the decision to expose the underlying platform mechanism. In retrospect, this was probably not the right decision. These methods are significantly different.
The JVM String.compareTo is an ordinal comparison of UTF-16 strings. This is a straightforward lexicographic comparison, performed character-by-character. The CLR String.CompareTo is culture-sensitive, i.e., the result of comparing two strings varies depending on the culture in effect at that time, which is thread-dependent.
There are several consequences of using CompareTo on the CLR, of varying import.
ClojureJVM and ClojureCLR differ on things such as comparisons.
(compare"a\u00ADb""ab");; => non-zero meaning not equal (JVM)(compare"a\u00ADb""ab");; => zero, meaning equal (CLR)
(\u00AD is the soft-hyphen character.)
And, thus, sort order:
(sort["b""B""a""A"]);; => ("A" "B" "a" "b") (JVM)(sort["b""B""a""A"]);; => ("a" "A" "b" "B") (CLR)
The difference here is that compare uses String.CompareTo (culture-sensitive) and = uses String.Equals (ordinal comparison).
sorted-set and hash-set yield different sets on the same inputs.
(count(sorted-set"a\u00ADb""ab"));; => 1, one string silently dropped because it compares as equal (CLR)(count(hash-set"a\u00ADb""ab"));; => 2, because hash sets use ordinal = (CLR)
Culture-sensitivity bites in some related places where string comparisons occur.
Culture-sensitive string compares are thread-dependent.
For example, ASP.NET Core sets the culture per request from Accept-Language, so (sort names) silently follows each user’s collation. Picking up information from the thread is fine; I feel it is better for it to be a deliberate choice rather than delivered behind your back.
Culture-sensitive comparisons are more expensive than ordinal comparisons.
Sorting a bunch of strings or creating a sorted map under ordinal comparison executes roughly 54-61% fewer instructions (instruction counts measured on one machine, not timings) than a culture-sensitive comparison using en-US. Capturing one culture comparer at startup to avoid thread lookup of the culture decreases the instruction count by roughly 5% compared to String.CompareTo today.
Culture-sensitive comparisons have not been consistent over time.
“Before .NET 5, the .NET globalization APIs used different underlying libraries on different platforms. … If you upgrade your app to target .NET 5 or later, you might see changes in your app even if you don’t realize you’re using globalization facilities.” (See Globalization and ICU - .NET). And you can switch globalization providers. ClojureCLR on .NET Framework uses NLS, not ICU, so Framework and .NET builds of ClojureCLR on the same machine can yield different results. That’s just on Windows. Let us not discuss Linux and Mac. Read it and weep.
Going deeper
The original focus in the benchmarking investigation surfaced the compare issue outlined above. After the initial results, I decided to expand the search to all string manipulation in the ClojureCLR, with comparisons against ClojureJVM where appropriate. Most of this internal string manipulation is related to reading data (which in Lisp-land includes programs) which arguably should not be culture/locale sensitive. It is not in ClojureJVM – all is ordinal. There are a few places where culture-sensitivity snuck into the ClojureCLR code, by carelessness or ignorance. Some are so marginal that I’m guessing they have never been encountered. For example, if we let <SHY> represent the soft-hyphen character, when reading source code:
JVM reads foo:<SHY> => creates symbol foo:<SHY>
CLR reads foo:<SHY> => throws an exception under en-US (at least)
The most significant culture-dependent bugs are
tr-TR: the flag to turn on direct linking in the compiler is ignored – the dotless ‘i’ (U+0131) makes an appearance.
sv-SE: (+ 1 1E-10M) doesn’t compile. Swedish uses a different minus sign character.
I consider these examples to be bugs. Fixing them is not a breaking change.
There are five functions (starts-with?, ends-with?, index-of, last-index-of, replace-first) in the clojure.string library that are in conflict with the JVM version and also demonstrably incorrect. These bugs are visible to users today, but they are bugs and should be fixed. Example: (replace-first "a<SHY>b" "ab" "X") gives "Xb" and corrupts the string. Results would change only for strings containing ignorable characters (soft hyphen, NUL, combining marks) and the changes would match the JVM’s answers.
A proposal
I plan to make two sets of changes. The first set is to fix the bugs above: the reader, the compiler, number literals, and clojure.string. I am not <SHY> about making these changes.
The second set of changes will be more user-visible and thus have the potential to break user code. This involves changing the definition of compare to be ordinal-based. To minimize impact, we can make this change selectable at startup. Set the switch to ‘culture’ and you get the current behavior; set it to ‘ordinal’ and you get ordinal-based comparisons. The ‘ordinal’ mode is pretty much identical to ClojureJVM behavior.
There is a third possibility: A captured culture mode which at startup creates a string comparator based on the CurrentCulture in effect at system startup. That is used by the compare function. If you change CurrentCulture, it will not be seen by this; there is no thread-dependency. You don’t get the large speedup, but do get the 5% savings from the no-thread-lookup. I’m not sure it is really worth it. The audience would seem to be a culture-mode user who sorts heavily and doesn’t care about threads. I’m not sure that’s much of a market. (This would require a little more testing before full validation as an option. )
For ordinal mode users, varying culture in things such as sorting is still possible. Most of the Clojure-defined functions, such as sort and sorted-set, have variations that take a comparator function. In fact, I think the general advice should be to be explicit in your sorting and comparator options always. To make this easier, we will add a culture-comparator function that will return a comparator function based either on the value of CurrentCulture when called or a supplied culture. You could write (sorted-set-by (culture-comparator "sv-SE")).
The choices
The bug fixes will happen regardless. For the compare changes, we have some choices.
Do nothing. Not really an option, from my viewpoint.
Put in the switch. Default is ‘culture’. If you want performance, you have to ask for it.
Put in the switch. Default is ‘ordinal’. Performance and consistency are the default.
To be clear, my own preference is #3. The current situation has internal inconsistencies, is inconsistent with the JVM, and is slower overall.
An additional choice is
Do we need the ‘captured culture’ mode?
I plan to put a link to this document up on the #clr channel in the Clojurians Slack for feedback. I’ll amend this post to reflect that discussion.
The result
TBD.
AI disclaimer/acknowledgement
I’ve been using AI coding tools extensively in my benchmarking work. There is a lot of tedious coding involved and the tools do it for me a lot faster and more correctly than I would on my own. This document reports on the results of that work. I drafted the document, then used the AI coding tool to vet it for accuracy, resulting in some typo corrections and a few suggested rewordings of inaccurate phrases. A stray comment in its review of my first draft sent me down a new line of inquiry that ended up in the two-phase approach here. The instructions to the agent include not drafting any code that might go into the ClojureCLR code base. It can identify locations and make prose suggestions only. (And it can write all the testing code its non-existent heart desires. It makes measurement runs on my commits as we go along.)
Just before this year's Clojure/conj, the Clojure team released a new
Clojure CLI REPL. Per the
README:A REPL for the Clojure CLI featuring multi-line editing with proper indentation, bracket highlighting, structural editing, inline eval, doc lookup for Clojure and Java, a data inspector, configurable prompts, keybindings, and much more.
In the age of generative models, the most important skill for a developer is to be able to recognize the shape of the problem and pick the correct way to express it. What's relevant today is the ability to do high level reasoning about algorithms, data structures, and data flows within the system. Imperative programming is quickly becoming akin to writing assembly because language models are quite competent at writing code in the small, while they stumble at high level design and architecture. So it is good to learn a language that operates at a higher level, such as Clojure, which is data-centric, composes functions declaratively, and keeps code close to the shape of the problem.
LLMs can generate code much quicker than I can, but the issue is how to test that the generated code does what I want. Before you can test anything in many languages, you have to recompile the program, and that may take quite a few minutes for larger projects. The length of that feedback cycle, in turn, sets the pace of your progress.
We also need to talk about the tedious reality of rebuilding application state. That matters even more when you are not writing the code yourself, so the output is inherently less intentional. Clojure collapses that loop because our workflow does not draw a hard line between when code is read, when it is compiled, and when it runs. Clojure runs in a live process, and you can redefine a function at the REPL and have the new version take effect immediately, without restarting. State stays in place while you change the code that operates on it. Inspecting the state to reproduce behaviors and verify fixes can be done instantly when you can reach into a running process.
An agent can similarly connect to a REPL to diagnose an issue and swap out the code without any downtime. Agents fundamentally need observability in order to get useful feedback about the changes they are making. Working with the REPL means the agent doesn’t have to go through all the steps of compiling and rebuilding the app, then logging its output to see the result. That translates into having to do fewer iterations to reach a working system, which becomes particularly valuable when working with a large codebase. The functionality of your system can keep evolving as you load new code into the running process without any restarts.
Another advantage comes from immutability, which helps reliably control the operating context of the program environment. When the majority of the logic in an application is written using pure functions, the agent can safely consider and test pieces in isolation without having to reason about the entire program. Agents can write functions one at a time, test them in the running program instead of making a whole bunch of changes, then running tests to find out if they worked.
At this point, I find anything with a compile cycle is a nonstarter if I have an option to have a live programming environment. It is not just compile and startup time that ends up being painful. A bigger problem is the tedious necessity of having to reconstruct the desired state every single run. When you have something small with limited functionality, that is fine, but as your application grows, rebuilding the state can take a significant effort. You might have to click through some menus in the user interface, wait for data to be processed from a service or a database, and so on. Being able to put your application in a particular state and make changes within that context is just a qualitatively better development experience.
Then there's the powerful macro system , allowing you to adapt the language to the problem domain, eliminating a lot of boilerplate code you would have to write otherwise. What makes this particularly powerful in Clojure is homoiconic syntax, where code is written using data structure literals. Since there is a common syntax for expressing both logic and data, a program can take any piece of code and manipulate it as it would with any other data structure, then evaluate it. This makes it incredibly easy to add new semantics because all you need is to make templates out of code. A macro works much like a function that accepts the code you wrote and produces the form that actually gets evaluated. The expression for printing a string does the printing when you evaluate it, but it is also nothing more than a list of the println symbol and the string itself.
One major advantage of S-expression based syntax is that it's easy for both humans and machines to read. And since the state can be trivially serialized as plain data structures, an agent can dump it in the REPL to inspect what’s happening in the application at any time. The data orientated nature of the language is a natural fit for LLMs since these models operate on text, making it exceptionally easy for them to see how data flows through the system without any opaque object graphs to worry about.
Furthermore, all the functions operate on a common set of data structures, allowing you to compose them together to transform data like Lego blocks. Clojure programs tend to be far more concise since the code is largely written through declarative composition of functions from the standard library, which encapsulate the implementation details. The result is that a codebase tends to be much shorter, leaving far less code to repeat. This conciseness matters since a smaller program costs fewer tokens, and fewer tokens leave more room in the context, making the language friendlier for local models.
In my experience, models lacking the broader context while making changes in a piece of code is one of the most common failure cases. Having a terse syntax means that more relevant code lives directly in the context window, directly addressing the problem. The model gets a significantly better view of what you are trying to do and can make much better decisions. If the whole call graph is sitting in the context window, the model sees how all the pieces fit together.
All these features combine to make a perfect environment for the agent to work in. Expressive syntax leads to less code repetition. Macros let you fold repeated patterns into new domain specific constructs. Code itself is just structured data that can be inspected and transformed. And the REPL ties it all together, providing you with a living system that evolves along with your code.
There are, however, a few drawbacks to the Java virtual machine, which Clojure traditionally runs on. From my experience, many developers, fairly or not, have an issue with requiring the JVM, and shy away from Clojure because of startup time, a somewhat heavyweight runtime, and perceived bootstrapping complexity.
One goal for Jolt in particular is to get more interest from outside the existing Clojure community by addressing these concerns. The compiler is a single binary, and it ships with all the tooling, such as dependency management and task running, baked in. It interops seamlessly with the native ecosystem via FFI, so you can use it in a way comparable to Python. Best of all, program distribution involves building a standalone binary similarly to Go. Thus, Jolt may remove the last big source of friction for trying Clojure.
In an age where writing code is cheap but verification and iteration remain expensive, the high level declarative style lines up exactly with what both agents and humans need to produce working code. Clojure is a great language to learn today because of the unique way it fits the era of language models.
Your database documentation says REPEATABLE READ. Your code assumes it. Have you ever checked?
I spent the last stretch building adya, a black-box checker for transactional isolation. You give it a log of transactions a database ran, and it tells you which isolation guarantees held. For each one that didn't, it prints the transactions involved and the chain of reads and writes that proves the violation.
It's an independent Rust implementation of the approach from Elle (Kingsbury and Alvaro, VLDB 2020), the checker behind Jepsen's database analyses. If you know Elle, adya reads the same history formats and takes the same flags. If you don't, keep reading.
The problem with "it passed the tests"
Isolation bugs don't show up in unit tests. They need two or more transactions to interleave in one particular way, and when they happen nothing crashes. You just get a balance that's off, or two people booked into the same seat.
The classic example is write skew. Two doctors are on call, and the rule says at least one must stay on call. Each doctor's transaction checks "is the other one still on call?", sees yes, and takes themselves off. Both commit. Now nobody is on call.
Snapshot isolation allows this. Postgres's REPEATABLE READ is snapshot isolation, so it allows this. If you thought REPEATABLE READ meant "safe enough", the bug ships.
Elle's insight was that you can catch this from the outside. Record what every client asked for and what it got back, infer the dependencies between transactions, and look for cycles that a given isolation level forbids. Atul Adya's 1999 thesis catalogued those cycles (G0, G1c, G-single, G2 and so on), which is where the name comes from.
Why another implementation
Elle is excellent. It's also a Clojure library on the JVM. People who want it outside a Jepsen test end up shelling out to elle-cli from a Go or Python harness. I found two open-source projects doing exactly that in CI (barn, bytecaskdb), and the bytecaskdb PR lists what went wrong: a blocked Clojars mirror, a crash when graphviz was missing, log lines corrupting the JSON output.
On my Windows machine, elle-cli 0.1.11 hung with no output on every anomalous history I gave it. On Linux it worked, but in my CI comparison it needed more than five minutes on five of forty random 300-transaction histories, and ran out of a 6 GB heap on one more. adya checked all forty in 0.2 seconds.
Elle also leaves the workload to you. You have to write the client that generates transactions, runs them against your database and records the history. That's the part most people never get to.
So adya ships both halves in one binary:
cargo install adya
# Run a workload against Postgres at REPEATABLE READ,# then check the history against serializability.
adya run postgres --url postgres://localhost/test -i repeatable-read -c serializable
What it looks like
You don't need a database to try it. adya has a built-in simulated database that implements isolation levels the textbook way:
$adya run sim -i snapshot-isolation -c serializable -n 300 -p 4 --seed 1
history.jsonl false
G2-item #0
Let:
T234 = {"index":234,"process":1,"type":"ok","value":[["r",10,[4]],["append",3,6],["r",3,[4,5,6]]]}
T237 = {"index":237,"process":3,"type":"ok","value":[["append",9,13],["append",10,5],["r",3,[4,5]],["r",10,[4,5]]]}
Then:
- T234 < T237, because T234 did not observe T237's append of 5 to 10.
- However, T237 < T234, because T237 did not observe T234's append of 6 to 3: a contradiction!
That's the doctors problem with list keys. T234 read key 10 before T237 appended to it, so T234 has to come first. T237 read key 3 before T234 appended to it, so T237 has to come first. Both can't be true, so no serial order exists. Snapshot isolation allows that cycle. Serializability doesn't.
How it works
The trick that makes this tractable is the workload. adya's default, borrowed from Elle, is list-append: every key holds a list, transactions append unique numbers to lists and read whole lists back. Each read then tells you the order of every append before it. If one client read [1, 2] and another read [1, 2, 5], you know 5 came after 2, without asking the database anything.
From those orders adya builds a dependency graph over transactions, using Adya's three edge types:
ww: T2 appended right after T1's append to the same key.
wr: T2 read a list ending in T1's append.
rw: T1 read a list that T2 later appended to, so T1 didn't see T2's write.
If you're checking a model with real-time guarantees (strict serializability), it adds edges for "T1 finished before T2 started", using a transitive reduction so the graph stays roughly linear in the history.
Then it looks for cycles, one strongly connected component at a time. Each anomaly class is a cycle with constraints. G-single has exactly one rw edge. G2-item has at least two, with two of them adjacent. G1c has no rw edges and at least one wr. Instead of enumerating cycles and classifying them afterwards, adya runs a breadth-first search over pairs of (transaction, path state). The path state is seven bits: how many rw edges so far (0, 1, or 2+), whether the last edge was rw, whether the first one was, whether two were adjacent, whether it has seen a wr, and whether it has used a real-time edge. One BFS then returns the shortest cycle of exactly the shape you asked for.
Before searching, it runs cheaper existence checks. For each model it asks whether the subgraph that model forbids cycles in has a nontrivial SCC at all. If the ww/wr-plus-one-rw subgraph is acyclic, snapshot isolation holds as far as cycles go, and adya skips every search that could only find SI violations. It also searches the most severe anomalies first and skips anything they imply. A 100,000-transaction history (30 MB of JSON) checks in about 1.6 seconds on my laptop.
Is it right?
A checker that's wrong is worse than none, so this took most of the work.
Elle's own expected results. elle-cli's test suite ships 56 list-append and rw-register histories along with the JSON verdicts Elle produced for them. adya matches all 56: same verdict, same anomaly types, and the same weakest models ruled out. Getting there taught me two Elle details I would never have guessed. It labels an edge that carries several relations by a fixed priority (ww before wr before rw before real-time), but tests whether a cycle exists using any relation the edge carries. And it counts a composed edge that loops back to the same transaction as a cycle.
Elle itself, live. On every push, CI generates random histories with adya's simulator and checks each with both tools. In the first full run Elle finished 34 of 40, and adya agreed on all 34.
Databases with known answers. The simulator implements serializable, snapshot isolation, read committed (with write locks), read uncommitted, and a deliberately broken "snapshot" that writes back stale state. Tests assert that correct runs produce zero anomalies at their own level and the expected ones above it. That test caught a bug in my simulator, not in the checker: my first read committed took no write locks, which is weaker than any real database.
What Postgres and MySQL did
CI runs 4,000 transactions from 10 clients over 6 hot keys against Postgres 17 and MySQL 8.4, at each isolation level, and checks every history against a ladder of models. List-append results:
database, level
serializable
snapshot isolation
read committed
Postgres READ COMMITTED
G-single, G2-item, internal, lost update
G-single, internal, lost update
valid
Postgres REPEATABLE READ
G2-item
valid
valid
Postgres SERIALIZABLE
valid
valid
valid
MySQL REPEATABLE READ
G-single, G2-item, internal, lost update
G-single, internal, lost update
valid
MySQL SERIALIZABLE
valid
valid
valid
Postgres did what its docs say. REPEATABLE READ is snapshot isolation, write skew included, and SERIALIZABLE came out strict serializable. I also restarted Postgres every few seconds during a 6,000-transaction SERIALIZABLE run (--fault "docker restart -t 0 pg"). Nine restarts, 2,957 commits, 3,043 failures, and the history still checked clean.
MySQL's REPEATABLE READ is weaker than its name. It lets a transaction lose another's update and see part of another transaction's writes, so it isn't snapshot isolation. Jepsen reported the same thing about MySQL 8.0.34 in 2023. adya reproduces it from a cold start in a few seconds.
Testing your own database
The built-in drivers cover SQLite, Postgres and MySQL. For anything else there's adya run exec: adya starts your client once per process, writes one JSON line per transaction to its stdin, and reads back one line saying what happened.
The repo has a 59-line Python client for SQLite as a template. Swap the SQL and you're testing your own database.
Limits
adya covers list-append and read-write register workloads. It doesn't do predicate reads, or Elle's bank, set and counter checkers. It prints text proofs instead of Graphviz plots. A clean result means it found no anomaly in that history. It doesn't prove your database is correct, so run it longer, with fewer keys for more contention, and with faults.
If you run it against something interesting, I'd like to hear what it found. Issues and PRs are open.
I use coding agents for basically all of my day day to work now. Recently I’ve been seeing more and more consumer-facing agent-like products and have been trying them out but find none are really able to do the things I want, mostly because they are all run by sketchy tech megacorps that I don’t trust and don’t want to connect my data to. Unfortunately Siri still really sucks, but in theory it’s more in the realm of what I actually want and Apple already owns my entire digital life anyway.
It made me wonder what a useful personal agent would be like. These are at least some of the requirements:
It needs access to all of the things I already use. For me that means protonmail for email, apple calendar and reminders, my obsidian personal wiki, and more. I’m not migrating my digital life to accommodate a bot.
It needs to remember things between sessions. LLMs are stateless and that makes them bad at long running tasks without harnesses that manage context.
It needs to be able to carry out long running tasks, like researching things for me or doing grunge work like organizing all my stuff. Getting a handle on the total dumpster fire of notes and todos I have scattered across a dozen different tools would be legitimately useful to me.
It needs to be able to schedule things for itself. Not everything needs to be a reminder visible to me. Some things I want are like “check if this computer is on sale yet”, “check for discounts on flights”. There are lots of little tasks I do throughout the day that are like this but I can’t be bothered to script them.
It needs to live in one place but be accessible anywhere, like on my computer and phone at least. Probably I’ll want to use it from multiple computers.
It would be really cool if I could even share them. There are some projects like volunteer orgs I’m a part of or home renovations that I collaborate with other people on.
It should at least be able to collaborate with other agents.
It should be self-managing and self-improving. I should never have to explain something twice. Over time it would just absorb my preferences and accumulate tribal knowledge about my life and projects, like a good assistant.
It needs to know how to use or maybe even make apps. I hate chat as a human-computer interface, it’s too unspecific and slow for most of what I want to do with computers.
Technically I think these requirements imply at least some of:
It needs to be able to read and write files.
It needs to be installed on one computer that never sleeps and be accessible to others over my tailnet, or similar.
It needs access to the internet though and probably needs a browser.
At least some of the agents need access to my actual computer. I don’t think there’s a practical way to give them access to my Apple or protonmail accounts but they could do everything they need from my personal Mac directly.
Anyway there are probably a lot more things that will come up but thinking about this has made me realize I want to try to build this. We’ll see how it goes!
Welcome to the first edition of The Pull, our new monthly Datomic newsletter rounding up projects and announcements from the Datomic community and the Datomic team.
Community Projects
EACL v8 for Datomic Pro is now available on Clojars as RC1. EACL (Enterprise Access ControL) is a situated ReBAC authorization library inspired by SpiceDB, built in Clojure and backed by Datomic Pro, Datahike, Datalevin or DataScript, though Datomic Pro remains the primary backend. Read the announcement or try the demo.
Trydatomic.org, an interactive website to learn how to query a Datomic database using Datalog, just got a content and design refresh.
Datomic Team News
In case you missed it: In April we released Datomic 1.0.7622, a big, feature-filled changelog with several performance improvements. It includes:
Feature: Read-only connections to storage and backups
Performance: Reduce CPU and memory required to calculate index metrics
Fix: Regression introduced in 1.0.7556 which broke the ddb-local protocol
Also earlier this year, Joe Lane, principal engineer at Nubank on the Datomic Core Dev team, gave a talk at the Unlocked Conference about Immutability in Motion, which explores how immutability and multi-tier caching power database performance at massive scale.
Recently Nubank engineers João Nascimento Mello and Mateus Oliveira, supported by engineers Gabrielle Cadurim and Carolina Silva, presented an online Day of Datomic workshop in connection with the 2026 Clojure Conj, now available to watch on YouTube.
We’re hosting DatomicConf this December 11, 2026 in Durham, North Carolina. Registration and CFP are now open, so join us there. Details at conf.datomic.com.