vim-slime

<!DOCTYPE html> <html> <head> <title>Tarn Barford</title> <meta charset="utf-8"/> <link rel="icon" type="image/x-icon" href="/favicon.ico"> <link href="/style.css" media="screen" rel="stylesheet" type="text/css" /> <link rel="alternate" type="application/atom+xml" title="Journals of Tarn Barford" href="/atom" /> <link href="/highlight.css" media="screen" rel="stylesheet" type="text/css" /> <link href="/highlight-console.css" media="screen" rel="stylesheet" type="text/css" /> </head> <body> <div id="container"> <div id="header"> <div id="header"> <p>From the <a href="/journal">Journals</a> of <a href="/">Tarn Barford</a></p> <h1> vim-slime </h1> <p> Mar 26, 2012 </p> </div> </div> <div id="post_content"> <html><body><p>Today I found the awesomeness that is <a href="https://github.com/jpalardy/vim-slime">vim-slime</a>, it's been an exciting day for me. <a href="http://common-lisp.net/project/slime/">Slime</a> is the "The Superior Lisp Interaction Mode for Emacs", I can almost hear the emacs crowd laughing.</p> <p>For those that use vim and haven't used Slime, vim-slime or <a href="https://github.com/vim-scripts/VimClojure">something similar</a>, this is why it's awesome:</p> <p><strong>Text can be sent from any process to the stdin of a <a href="http://www.gnu.org/software/screen/">gnu screen</a> or <a href="http://tmux.sourceforge.net/">tmux</a> session. The process in this case is vim and the screen/tmux session is a terminal</strong>.</p> <p>Screen is a <a href="/journal/oh-screen-where-have-you-been">really neat</a> terminal multiplexer (you can run multiple terminals in a terminal window). The multiplexed shell processes are children of the screen process, which itself is not a child of the terminal window process. This means a screen process and its child processes keep running if you close the terminal window. Later you can re-connect to it, this is what makes vim-slime possible.</p> <p>Here is an screen shot, on the left is me in gVim writing some awful Clojure <a href="#footnote-1">[1]</a>. On the right is a screen buffer in which I started a Clojure REPL. When I want to try run some code I can send any vim text selection to the REPL in a keystroke (or two).</p> <p><img alt="vim slime screenshot" src="screenshot.jpg"/></p> <p>It doesn't have to be a Clojure REPL either, we can send anything to a screen shell. We could run git commands, find, grep, sed, etc. Like with the Clojure REPL we can even interact with any terminal programs that use STDIN.</p> <p>This concept can be taken even further, You can even connect to a tmux session over SSH and share a terminal or a <a href="http://remotepairprogramming.com/remote-pair-programming-with-tmux-and-vim-the">terminal program like vim to do remote pairing</a>!</p> <p>Hopefully remote pairing is the topic of my next post as there are a couple geographically distant people I know who are keen to do some pair hacking. I stand to learn a lot!</p> <p><a name="footnote-1">[1]</a> I learnt almost everything I know about Lisp from <a href="http://www.ccs.neu.edu/home/matthias/BTLS/">The Little Schemer</a>. Great book.</p></body></html> </div> <div id="comments"> </div> </div> <div id="footer"> <p>&nbsp;</p> <p>Questions, comments, suggestions? Email me, <a href="mailto:tarn@tarnbarford.net">tarn@tarnbarford.net</a> (<a href="/pgp.txt">public key</a>)</p> <p>&nbsp;</p> </div> </body> </html>

Permalink

Swipe Keyboard

<!DOCTYPE html> <html> <head> <title>Tarn Barford</title> <meta charset="utf-8"/> <link rel="icon" type="image/x-icon" href="/favicon.ico"> <link href="/style.css" media="screen" rel="stylesheet" type="text/css" /> <link rel="alternate" type="application/atom+xml" title="Journals of Tarn Barford" href="/atom" /> <link href="/highlight.css" media="screen" rel="stylesheet" type="text/css" /> <link href="/highlight-console.css" media="screen" rel="stylesheet" type="text/css" /> <style> #swipe-canvas { position: relative; width: 900px; height: 300px; } #swipe-results { font-size: 30px; padding-left: 50px; padding-left: 50px; } #swipe-results ul { margin: 0px; padding: 0px; } #swipe-results li { float: left; background-color: #DDDDDD; list-style-type: none; padding: 10px; margin: 5px; border-radius: 5px; } #swipe { position: relative; } #swipe-loading { position: absolute; height: 50px; width: 300px; top: 85px; left: 300px; background-color: darkgray; border-radius: 10px; text-align: center; padding-top: 20px; border: black; border-width: 5px; } </style> </head> <body> <div id="container"> <div id="header"> <div id="header"> <p>From the <a href="/journal">Journals</a> of <a href="/">Tarn Barford</a></p> <h1> Swipe Keyboard </h1> <p> Apr 06, 2014 </p> </div> </div> <div id="post_content"> <html><body><p>When I first tried a <a href="http://www.swype.com/">Swype</a> keyboard I was impressed how effective it was. Even though I don't use the feature on my phone I was interested in how it could be built, so I <a href="https://github.com/tarnacious/swipe-keyboard">implemented this otherwise useless swipe-able keyboard</a> below. It probably doesn't work on mobile devices, but works on modern browsers with mouse pointers (although I've only really tried Chrome and Firefox).</p> <div id="swipe"> <canvas height="300px" id="swipe-canvas" width="900px"></canvas> <div id="swipe-results"></div> <div style="clear: both"></div> <h2 id="swipe-loading">Loading<noscript>Javascript is Required</noscript></h2> </div> <p>I initially tried to solve this using the technique Peter Norvig famously uses in his <a href="http://norvig.com/spell-correct.html]">spell checker</a>. He takes a sequence of characters and generates a set of word candidates by adding, removing and swapping characters in the original sequence, the generated candidates are removed if they are not found a dictionary. This can work but to be effective too many combinations need to be generated.</p> <p>If the dictionary is indexed into a <a href="http://en.wikipedia.org/wiki/Trie">trie</a> the number of combinations generated can be reduced significantly by traversing the trie and only generating valid letter combinations. This is a pretty bare implementation of that, it requires: </p> <ul> <li>The first and last characters of the initial sequence are used </li> <li>Intermediate characters in the initial sequence can be repeated or discarded </li> <li>No characters are added or swapped</li> </ul> <p>Basically, if you swipe through all the characters in a word in order, then the word will be found if it is in the index regardless how many characters are swiped in between. It is surprisingly quick and effective.</p> <p>This implementation uses <a href="https://raw.github.com/first20hours/google-10000-english/master/google-10000-english.txt">these 10000 words</a>, I intended to use digital books but never got around to it as these words demonstrate the concept well enough.</p> <p>This is the first thing I've written in <a href="https://github.com/clojure/clojurescript">ClojureScript</a> or <a href="https://github.com/clojure/clojurescript">Clojure</a>, so my code my vary from non-idiomatic to shamblolic. I initially used a <a href="http://clojuredocs.org/clojure_core/clojure.zip/zipper">zipper</a> to build the trie with immutable data structures, but found the indexing took to long with my zipper implementation so I <a href="https://github.com/tarnacious/swipe-keyboard/commit/6edd7b26e78121fbe8586b3f0ef54ca8277d9e32">switched to using native Javascript maps</a>.</p> <p>I found that <a href="https://github.com/clojure/core.async">core.async</a> library is really awesome, the <a href="http://docs.closure-library.googlecode.com/git/index.html">Google closure library</a> and <a href="https://developers.google.com/closure/compiler/">compiler</a> integration with <a href="http://leiningen.org/">Leiningen</a> the <a href="https://github.com/emezeske/lein-cljsbuild">cljsbuild plug-in</a> to be impressive. My main pains were the slow JVM start-up time, the advanced closure compiler build of the web worker script fails silently when run (but the main script works fine when compiled with the advanced compiler), and at times I felt some compile time type checking would be nice.</p> <p>I would like to extend this experiment to index the word occurrence counts and proceeding word counts in original text and rank the found words as most likely. Support casing, umlauts, special characters, spelling correction and compound words in the indexing and lookup. I think a live lookup while swiping would also be possible.</p> <p>Overall this was fun, turned out OK I think, and was a great learning experience.</p></body></html> </div> <div id="comments"> </div> </div> <div id="footer"> <p>&nbsp;</p> <p>Questions, comments, suggestions? Email me, <a href="mailto:tarn@tarnbarford.net">tarn@tarnbarford.net</a> (<a href="/pgp.txt">public key</a>)</p> <p>&nbsp;</p> </div> <script src="swipe.js" type="text/javascript"></script> </body> </html>

Permalink

Measuring agent workflows on public benchmarks

An agent that books invoices, answers questions over a database or fills in spreadsheets is more than a language model. It is a workflow: the model, the instructions it is given, the tools it may use and the steps it takes. Every part of it can be changed, and every change has a price. Measuring a workflow on a fixed set of cases is the statistical counterpart of a regression test, and it is what makes a change safe to deploy. This article shows how Dvergr, our open-source system for running and measuring agents, measures such workflows, what it found when we measured ten models on two public benchmarks, and why the most useful thing a measurement finds is often not the best model but a better workflow.

The question behind a leaderboard

Suppose a firm wants an agent to answer its staff’s questions about the company database. It has to choose a model, write the instructions, and decide what the agent may do: write one query and stop, or look at the data first, try a query, read the error and try again. A public leaderboard answers a narrower question, which model scored highest on someone else’s cases, with someone else’s instructions, on one run.

Dvergr answers the firm’s question directly. It takes a set of cases, each a task with the answer a person recorded for it, and a set of candidates, each a model together with its instructions and tools. It runs every candidate on every case and reports, for each candidate:

  • how often it is right, with a range that says how much that number could move on another set of cases of the same kind;
  • where it fails, check by check;
  • what a correct answer costs at the model’s list price;
  • how long it takes.

A firm that deploys an agent needs what software teams get from regression tests: a fixed set of cases with recorded answers, run again on every change. An agent’s output varies from run to run, and a provider can change the model behind a name it keeps, so one passing run says little. The result of a run is therefore a rate with an interval rather than pass or fail, and a change is judged by comparing it with the previous version on the cases where the two disagree.

We tried this on two public benchmarks, where anyone can check the cases and the answers. SpreadsheetBench asks for changes to Excel workbooks. BIRD asks questions in plain language about real databases, to be answered with a query.

The models come in families, named here from cheapest to most expensive: OpenAI’s Luna, Sol and Astra, in generations 5.6 and 6 (6.1 for Sol); Anthropic’s Claude Haiku, Sonnet and Opus 5.5; and two models with published weights, GLM and DeepSeek.

Results at a glance

BIRD, 100 questions, SQL SpreadsheetBench, 59 tasks 406080100 406080100 $0.001$0.01$0.1 $0.001$0.01$0.1$1 cost per correct answer (log scale) cost per passing task (log scale) % correct Luna 5.6HaikuDeepSeekGLMLuna 6Sol 6.1Astra 6 Luna 5.6Sol 5.6OpusHaikuSonnetLuna 6Sol 6.1Astra 6 GPT-5.6 GPT-6 Claude 5.5 open weights (Fireworks) Correctness with its 95 % interval against cost per correct answer, at list price, measured 2026-10-01 to 2026-10-10. BIRD: seven models writing SQL for SQLite on the same 100 questions. SpreadsheetBench: eight models on the same 59 tasks. Overlapping intervals are ties. On BIRD, GPT-6.1 Sol and GPT-6 Astra score highest, the first models we measured to rise above the rest, though their intervals still overlap the others'; on SpreadsheetBench the models tie, at costs that differ by more than a hundredfold.

Each dot is one model. Its bar shows how far the score could move if we had picked a different 100 questions of the same kind (59 tasks on SpreadsheetBench), the way a poll of 100 people would come out differently with another 100. Each model answered each question once, and the bar follows from that count alone, which is why it is nearly the same length for every model: with 100 questions and scores around two thirds, a score can move by about nine points either way. Where two bars overlap, the measurement cannot tell the two models apart, and the choice between them comes down to cost and time, on the horizontal axis.

Because every model answered the same questions, two models can also be compared more sharply. Most questions do not help, because both models got them right or both got them wrong. What decides is the questions where they differ. If model A was right and B wrong on twelve questions, and the reverse happened on only two, A is better even though their bars overlap; six against five says nothing. We give these counts with a p-value, the chance that two equally good models would split their differences at least this unevenly, and below 0.05 we treat a difference as real. Dvergr’s report gives the bars; the counts and p-values here we computed from the per-question results it records.

The tables also give the median time per case. Time depends on how busy the provider was that day, so it compares candidates run together better than candidates run on different days.

How a measurement runs

benchmark cases one row per case + recorded outcome certify ungradable cases excluded, with reasons experiment room every candidate on every case report intervals, failures, cost and time fork A × case 1 fork B × case 1 fork A × case 2 … checker verdict, or fault to run again A measurement. Certified cases enter an experiment room; each attempt runs in its own fork and is graded there; only the verdict is kept.

A benchmark enters Dvergr as a case pack: the cases, their recorded answers, and a checker that decides whether an answer is right. Before any model runs, the pack is certified. Certification lists every case that cannot grade an answer and says why: an id that is missing or appears twice, an empty answer, the same inputs recorded with different answers, or a recorded answer the checker itself does not accept. Those cases are set aside. On SpreadsheetBench, certification kept 390 of the 400 tasks without spending any model tokens.

The measurement itself takes place in a room, Dvergr’s unit of shared state: its files, its database and its history. For each candidate and case, Dvergr makes a fork of the room, a private copy that costs almost nothing to make because it shares everything until something changes. The candidate works in its fork: the files and the workbook it changes belong to that copy, and a database it may only read is reached for each query through a fresh read-only connection or an unchangeable snapshot, so nothing it does reaches the original data or the candidates working beside it. The checker grades the attempt’s result (the files in the fork, the query it submitted, or its edits replayed on a fresh copy of the workbook), the verdict is recorded in the room, and the fork is thrown away. Dvergr’s own share of each case, the fork included, is a fraction of a second, against the ten to seventy seconds a model takes to answer.

The statistics depend on these forks. Because every candidate starts from the same state, two candidates’ answers to a case can be compared directly, which is what the paired comparison needs. Because no attempt can change what another sees, a candidate is graded on its own work. And because a fork shares everything that has not changed, running every candidate on every case stays cheap. This follows from how Dvergr keeps state: the room’s files and its database are forked with copy-on-write branches across git and Datahike, so a fork records only what it changes and never writes to the original.

Four rules keep the numbers honest:

  • The same cases for everyone. Every candidate answers every case, which is what makes the paired comparison possible.
  • A broken connection is not a wrong answer. When the provider is down or a request times out, the case is run again. A model that answers wrongly, or gives no answer, gets its verdict.
  • List price. Models are billed per token, a piece of a word, read or written. Cost is what the same tokens would cost through the provider’s public price list, whether a subscription or the firm’s own key paid for them, with text the provider has already seen recently (cached input) at its lower rate.
  • A changed setup is a new experiment. The cases, candidates, instructions, checker and software versions together define an experiment. An interrupted run resumes where it stopped, but if any of them has changed, running it again starts a new experiment instead of mixing two.

Does the agent need to explore?

The candidates above are models inside Dvergr’s own workflows. On BIRD, the model may run queries in SQL, the standard database query language, see their results or errors, and submit an answer when it is satisfied, within twenty turns. On spreadsheets it may read cells, write values and formulas, see what they compute, and submit.

A firm should know how much of a score is the model and how much is the workflow around it. An agent that explores uses more tokens than one that answers at once, and it is worth knowing what those tokens buy. So we also ran each benchmark’s own reference setup, with the same model on the same cases. These runs are not Dvergr workflows: they use each benchmark’s published code to build the prompts and grade the answers, and Dvergr only connects them to the model, so that both setups reach it the same way. The scripts, with the steps to reproduce each number below, are in the repository under benchmarks/reference/.

On BIRD, one instruction did most of the work

BIRD’s reference setup is a single prompt: the list of the database’s tables and columns (its schema), the question and a hint, and a request to write the SQL query. The model answers once, without seeing any data. We ran it as published, then with one line added, return exactly the columns the question asks for. Graded with BIRD’s own evaluation script, on the same 100 questions:

WorkflowGPT-5.6 LunaClaude Haiku 5.5Tokens per question (Luna)
BIRD’s prompt, one reply51531,200
the same, plus the line about columns61641,200
Dvergr: run queries, see results and errors, then submit65698,200

Two questions show what each part contributes. Question 781 asks for the heights of the heroes whose eye colours are amber. With BIRD’s prompt, Luna returned each hero’s name beside the height. The heights were right, but BIRD compares the returned table as a whole, and a table with an extra column is wrong. With the line about columns it returned the heights alone and passed. Of the fifteen questions Luna answered on our workflow and missed with BIRD’s prompt, seven failed for no other reason than such an extra column, and the one line is worth ten points for both models, a difference too large for chance (p = 0.013 and 0.003).

Question 817 asks for the race of the blue-haired male superhero, and its hint says the colour is written 'blue' and the gender 'male'. In the database they are written 'Blue' and 'Male', and SQLite compares text exactly. The one-reply query followed the hint and found nothing. In Dvergr’s workflow Luna wrote a query that ignores capitals, ran it, saw three heroes, and submitted the query for their races. That is the kind of question where looking at the data pays.

Exploring the data adds another four or five points. With 100 questions that could be chance (p = 0.34 and 0.30), and it costs seven times the tokens. For these questions the cheapest good workflow is a single reply with the right instruction, at about a third of the cost per correct answer. The agent that explores scores higher, but 100 questions cannot show that the gain is real, and it costs more.

We had written that line into our own workflow from the start, because we knew how BIRD grades. Until we measured it, we did not know it was most of what our workflow added.

On spreadsheets, the recalculating program decided the score

On SpreadsheetBench the scores depended less on the model than on how the answers were checked.

The benchmark’s reference setup shows the model the instruction and the first five rows of each sheet, and asks for a short Python program that edits the workbook. Its multi-round version runs the program and shows the model the output or error, for up to five turns. We ran both versions with GPT-5.6 Luna on our 59 tasks, each program in a Docker container that can see only its own input file, and graded them with the benchmark’s own comparison code. Dvergr’s workflow writes no Python: its agent edits the workbook directly, through a spreadsheet engine in Dvergr’s sandbox.

SpreadsheetBench’s grader reads each answer cell’s stored value, the result a spreadsheet program saved with the formula, and never computes a formula itself. A file written from Python stores formulas without values, so the benchmark’s instructions add a step before grading: open every workbook in LibreOffice (on Linux and macOS) or Excel (on Windows) and let it recalculate. We followed the instructions, using LibreOffice on Linux.

One task shows what happened. Task 56563 has a column of monthly amounts, one of which is an error value, and asks for a formula that adds them up anyway. Both workflows wrote the same correct formula into the answer cell, =AGGREGATE(9,6,C2:C13), which sums the range and skips cells holding errors. Its value is 77,772.6. We graded that one cell two ways:

Graded byWhat it found in the cellVerdict
the grader after LibreOffice recalculated, as the benchmark’s instructions say#NAME?: LibreOffice recognises functions added to Excel since 2010, such as AGGREGATE, only when the file marks them with a prefix (_xlfn.AGGREGATE), and neither workflow wrote itwrong
Dvergr’s spreadsheet engine77,772.6, as does LibreOffice once the prefix is writtenright

The same answer was wrong or right, depending on which program computed it. Across the 59 tasks:

WorkflowRecalculated by LibreOffice (the benchmark’s instructions)Tokens per task (median)Median time
SpreadsheetBench, one reply with code342,20015 s
SpreadsheetBench, up to five turns with execution414,30017 s
Dvergr: read and write cells, see what they compute, submit41 as written, 48 as Excel stores them5,70024 s

Graded the benchmark’s way, after LibreOffice’s recalculation, the multi-round workflow scores 41, and Dvergr’s workbooks as written also score 41; each is right where the other is wrong on eight tasks, so on that footing they tie. Written with Excel’s prefix, Dvergr’s workbooks score 48. Graded by Rechentafel, the open-source spreadsheet engine Dvergr’s agent works in, they score 51. Grading with our own engine would prove little if the engine had not been checked against Excel. The benchmark’s recorded answers are the values Excel saved, and for 390 of the 400 tasks Rechentafel recalculates each recorded answer to exactly that value, which is how certification works on this benchmark. The workflows differ less than the graders do. A firm measuring a spreadsheet workflow should grade with the program its people use, or with one checked against it.

Improving a workflow by measuring it

The line about columns is the kind of change a measurement should find, not one an author should have to know in advance. Dvergr is built so that variants of a workflow can be written and measured like any other candidate.

A workflow in Dvergr lives in its room as ordinary files: the task and its instructions, the checker, the cases, and any code the agent runs in its sandbox. Because they are files in the room, an agent working there can read them and write a variant: a sentence added to the instructions, a look at the data before the answer, a different model. Each variant is a candidate like any other, measured in its own forks on a set of tuning cases. A variant that does better is then measured on cases it has not seen, so that one tuned to the quirks of one set of questions does not pass for an improvement. A variant’s checker is trusted only after the host’s administrator promotes it, through an administrator’s connection. The promotion is recorded outside the room’s files, where no agent can write, so an agent cannot raise the trust of its own checker.

We took one step of this loop by hand on our Datalog workflow for BIRD. Datalog is the query language of Datahike, the database Dvergr keeps its own state in, and the workflow answers BIRD’s questions with it instead of SQL. After reading the failures of one run we added two sentences to its description of the language, and measured the change on another 100 questions that run had not seen:

WorkflowCorrect (of 100)Input tokens per question
SQL on SQLite598,880
Datalog, revised description6314,042
Datalog, previous description5816,684

The revision answers five more questions with less input, but at 100 questions that could be chance (p = 0.23). That is the honest result of one step of tuning: a promising change, to be confirmed on more cases before it is adopted. Today a person or an agent proposes each variant and Dvergr measures it. Letting an agent in the room run the whole loop, propose, measure and hand the winner to the owner, is what we are building next. The rule for adopting a variant is the one a team applies to a code change: merge it when the cases show it is not worse.

SpreadsheetBench in detail

This section and the next give the full results behind the chart, model by model. These numbers use our workflow and our spreadsheet engine, on 59 of the 390 certified tasks, so they are not comparable to the public leaderboard, where agents work with other tools.

ModelTasksPassed95 % intervalCost per passMedian time
GPT-5.6 Luna595176–93 %$0.002624 s
GPT-5.6 Sol595278–95 %$0.03327 s
Claude Opus 5.5 (Claude Code subscription)595278–95 %$0.06217 s
Claude Haiku 5.5 (Claude Code subscription)594768–88 %$0.004622 s
Claude Sonnet 5.5 (Claude Code subscription)594666–87 %$0.3319 s
GPT-6 Luna594870–90 %$0.000810 s
GPT-6.1 Sol595585–98 %$0.0109 s
GPT-6 Astra595585–98 %$0.0509 s

Paired with GPT-5.6 Luna, none of the GPT-5.6 and Claude models is measurably different. Opus is right where Luna is wrong six times and the reverse five times; Sol three and two; Haiku four and eight; Sonnet three and nine (the closest, p = 0.15). What differs is the cost. Per passing task Haiku costs about twice as much as Luna, Sol thirteen times, Opus twenty-four times and Sonnet over a hundred times, because Sonnet sometimes wrote very long answers: its slowest tenth of tasks took over nine minutes.

The GPT-6 models, run on the same tasks a few days later, are the first to move the top. GPT-6.1 Sol and GPT-6 Astra pass 55 each and disagree on only four tasks, two each way, so Astra costs five times as much as Sol for the same answers. Sol 6.1 is measurably better than GPT-6 Luna (eight tasks to one, p = 0.039) but not than GPT-5.6 Luna (five to one, p = 0.22). They were also about three times faster, partly because of the models and partly because of the day.

Every task but one was solved by at least one model, so most failures belong to a model, not to the task.

BIRD in detail

BIRD’s questions cover 11 databases. We can answer each question three ways: SQL on SQLite, the database BIRD ships with; SQL on Datahike, through its PostgreSQL-compatible layer pg-datahike; and Datalog on Datahike. Even before any model runs, 1372 of BIRD’s 1534 reference queries return exactly SQLite’s rows on pg-datahike.

With GPT-5.6 Luna on 100 fresh questions held out from tuning:

EngineCorrect (of 100)95 % intervalCost per correct answerInput tokens per questionMedian / slowest tenth
SQL on SQLite6555–74 %$0.00237,83910 s / 23 s
SQL on Datahike (pg-datahike)6555–74 %$0.002610,13610 s / 41 s
Datalog on Datahike6454–73 %$0.004415,06113 s / 42 s

The three engines answer equally well: each pair disagrees on seven to ten questions, split almost evenly. They differ in cost. In Datalog the model looks at the schema more before it answers, so a correct answer costs about twice as much as in SQL. Which query language to give an agent is a cost decision, and the measurement says how large it is.

Seven models on the same 100 questions:

ModelSQL on SQLite (of 100)Datalog (of 100)Cost per correct answer, SQL / DatalogMedian time, SQL
GPT-5.6 Luna6564$0.0023 / $0.004410 s
Claude Haiku 5.56968$0.0026 / $0.00479 s
DeepSeek V4.1 Flash (open weights)6664$0.0074 / $0.01958 s
GLM 5.3 Flash (open weights)6458$0.0042 / $0.01066 s
GPT-6 Luna6860$0.0009 / $0.00145 s
GPT-6.1 Sol7471$0.015 / $0.0198 s
GPT-6 Astra7570$0.067 / $0.0916 s

Among the first four models, no two are measurably different. Any two disagree on 10 to 18 questions; the most lopsided pair (GLM’s Datalog against Haiku’s, p = 0.013) is no more lopsided than chance alone produces somewhere among 28 comparisons.

GPT-6.1 Sol and GPT-6 Astra answer 74 and 75 in SQL, against 65 for GPT-5.6 Luna (eleven questions to two and twelve to two), and tie with each other. Counting all the comparisons we made, even that is not quite conclusive, but it is the first gain we have seen, and it comes within about five questions of the 80 that BIRD’s recorded answers allow (see below). Again Astra costs about five times as much as Sol per correct answer.

The two open-weight models (models anyone may download and run on their own machines), served by Fireworks, answer as many questions as GPT-5.6 Luna and Claude Haiku, at two to three times the cost per correct answer, because they used more tokens and none of their input was served from a cache. For a firm whose data may not leave its own machines, that is the useful result: a model it can run itself answers these questions as well, and its cost is then the firm’s hardware, not a list price.

What the benchmarks get wrong

A benchmark is only as good as its recorded answers and its grader, and both benchmarks have faults that change scores.

On BIRD, twenty of the 100 questions were answered correctly by none of the first four models in either language. On most of them the models agree with each other and not with the recorded answer:

Question (BIRD id)Recorded answerWhat the models answered
atoms of molecule TR346 and its bond types (309)a query for molecule TR000molecule TR346
the German type of a card (482)the English typethe German type
patients with abnormal CRP and no recorded data (1256)208, the number of lab rows25 patients; the only answer graded right, GPT-6 Luna’s in Datalog, repeated the mistake
expenses of the budget with the lowest remaining (1365)one row, cut off by LIMIT 1both expenses of that budget
difference of two percentages (1458)12.120.1212, following the hint’s formula, which has no ×100
eye colours of Marvel heroes by popularity (728)colour, count and a rank columncolour and count

We count about ten questions whose recorded answer is wrong, four more whose hint contradicts it, and the rest a matter of which columns or which wording. On these 100 questions the best attainable score is about 80, not 100. Paired comparisons survive this, since every candidate loses the same questions, but the differences shrink, and a model that answers the question as asked gets no credit for it. Our grader agrees with BIRD’s own on all 200 SQL answers we checked, so the scores above are the benchmark’s own.

On SpreadsheetBench:

  • The published evaluation script cannot score the published tasks. On the repository’s main branch it compares each task’s input workbook with the recorded answer (the line that reads the model’s output is commented out), and its file names are those of an earlier release. Every published score on the 400 verified tasks comes from a modified script that is not part of the benchmark.
  • Formulas are graded as empty unless the workbook is recalculated first, as described above. The benchmark’s instructions include that step, but the program that recalculates changes the score, and the step also recalculates the recorded answers, so after it they are that program’s results too.
  • One task can never pass. The answer position of task 45944 contains spaces, and the checker fails on it whatever the answer.
  • Some recorded answers are wrong or misfiled. Task 118-50’s answer misses a pair its own rule finds in the input, and the two strongest models were graded wrong for finding it. Task 42930’s answer file carries another task’s number.

Certification sets aside ten of the 400 tasks for reasons like these before any model runs.

Checking our own measurements

The setup that measures is code too, and it needs the same tests as the workflows it measures. Before publishing, four AI review agents read our benchmark code, documentation and stored results, each with one question: whether every benchmark runs the same way, whether the forks isolate each attempt, whether the documentation matches the code, and whether the numbers are right. They found nine problems, among them a provider outage counted as the model’s failure and cached input billed at its full price. Each fix came with a test that fails without it, and the affected runs were repeated or withdrawn.

What these numbers cannot tell you

Each model answered each case once, so a few answers could change on another run; the intervals say how much that matters. The numbers describe these cases with these workflows, and a different prompt, tool or model version can change them, which is why an experiment records all three. A tie at 59 tasks says what 59 tasks can separate, not that two models are equal. As with tests, coverage sets what can be found: with 100 cases, a regression of a few points can go unnoticed. And a public benchmark is not your workflow: it shows that the method works where others can check it, not how a model will do on your cases.

Where the method comes from

Seen this way, a workflow is a probabilistic program: a program whose output is drawn from a distribution, and measuring it is inference about how often that output is right. The Jeffreys interval in our reports is a Bayesian estimate of that rate. The view comes from probabilistic programming, which Christian Weilbach worked on in his doctoral research in Frank Wood’s group, with the Anglican language and the Daphne compiler (TMLR 2025). It continues in Foerster, our library for inference over forkable worlds, where alternatives are forks of the same state, as candidates are here. Dvergr’s measurements do not use Foerster yet. Two steps it would allow are stopping a comparison as soon as the evidence decides it and spending more cases where two candidates are close.

Try it on your own cases

The benchmark that matters for a firm is its own history: past bookings, coded invoices or resolved tickets, each with the outcome a person recorded. Dvergr is open source and runs on your own machine:

  1. Install and connect. Clone the repository (it needs the Clojure CLI and babashka), and add bin/dvergr-mcp --profile bench as an MCP server in Claude Code or Codex. The models it compares can be used through a Claude Code or Codex subscription or any API key; the README lists the providers, the setup and where Dvergr keeps its state.
  2. Turn your history into a case pack. Export a table of past cases as CSV, as Excel or DATEV write it, and name its columns: the id, the inputs, and the outcome with a rule for checking it (exact, a number within a tolerance, a date, a set). Certification reports which cases can grade an answer and why the others cannot, which is a first result in itself.
  3. Run the comparison. Ask your assistant to run catalog_benchmark with the models you want compared, or, from a shell, turn the CSV into a pack with clojure -M -m dvergr.catalog.casepack-cli and run clojure -M -m dvergr.catalog.room-run <pack> --models … --cases 30. The report gives each model’s success with its interval, where it fails, the cost per correct result and the time.

Running benchmarks and case packs in the documentation give the details. A following article applies this to a year of bank bookings exported from DATEV.

If you would rather we built the benchmark with you, we run fixed-fee pilots: six weeks on one workflow of yours, ending with the benchmark and a report on which model, instructions and steps to run, at what cost. Send us a note.

Method

ModelsGPT-5.6 Luna and Sol, GPT-6 Luna, GPT-6.1 Sol and GPT-6 Astra (Codex subscription); Claude Haiku 5.5, Sonnet 5.5 and Opus 5.5 (Claude Code subscription, each pinned to its exact version); GLM 5.3 Flash and DeepSeek V4.1 Flash (Fireworks). Each answers every case once.
CasesBIRD dev, the held-out third of each database’s questions, fresh samples of 100; SpreadsheetBench verified, held-out tasks among the 390 certified
GradingBIRD: the returned rows equal the recorded query’s rows; SpreadsheetBench: the answer cells equal the recorded workbook’s values
StatisticsJeffreys 95 % intervals for success rates, as Dvergr’s report gives them; exact McNemar test on the cases two candidates disagree on, computed from the recorded per-case verdicts
Costlist price per million tokens on the day of measurement, cached input at the cache rate
Reference setupsBIRD’s prompt from its published gpt_request.py and its evaluation script; SpreadsheetBench’s single- and multi-round drivers and its comparison code, the generated programs run in a Docker container; both reach the model through Dvergr’s model connection: benchmarks/reference/
Code and dataBIRD and SpreadsheetBench are public; Dvergr and both benchmark adapters are open source: github.com/replikativ/dvergr

Permalink

Week Notes 2026.41

Babashka tasks for deploying Datomic Cloud ions; making AWS API Gateway requests with SigV4 authentication.

Permalink

neorepl: A Terminal REPL for Clojure (and Friends)

For more than a decade the Clojure REPL most people saw in a terminal was REPLy, the one behind lein repl. It served us well, but it’s showing its age, and I kept wishing for a terminal REPL that felt like CIDER, without the Emacs part. So I finally built one. neorepl 0.1 is out today!

Why another REPL?

There’s no shortage of Clojure REPLs, so it’s fair to ask. Here’s what I was after:

  • nREPL all the way down. neorepl only speaks nREPL, so it works with anything that runs an nREPL server: Clojure, babashka, shadow-cljs, Basilisp, ClojureCLR and so on. The smarts live on the server. When cider-nrepl is there neorepl uses it, the same way CIDER does, and when it isn’t it falls back to nREPL’s built-in ops, and then to evaluating a bit of code.
  • A REPL that’s pleasant to use: a proper line editor, highlighting, completion, arglists as you type, pretty printing, history per project and a Ctrl-C that actually interrupts the evaluation.
  • A CLI that’s just as good for scripts and coding agents. Every lookup the REPL can do is also a command, with structured output and exit codes that mean something.
  • Something small and modern. Clojure 1.10+, nREPL and JLine 4 are the only dependencies, and the same code runs on the JVM and on babashka. The whole thing is about 3,300 lines right now, and I’d like to keep it that way.

rebel-readline and bb repl --connect are both great, and I borrowed from them freely, but neither tries to be an nREPL-first tool with CIDER’s tooling, jack-in and a CLI for agents. That combination is the whole point of neorepl.

For humans

Here’s the REPL in action:

The neorepl REPL

A few things worth pointing out:

  • Enter only submits the input once its forms are complete, so pasting a whole function just works.
  • TAB completes, with the candidates’ types next to them, and the arglist of the function you’re calling shows up after the cursor.
  • Values get pretty printed on the server and highlighted on the client.
  • Commands start with a comma, the CIDER way: ,doc, ,source, ,apropos, ,macroexpand, ,pst, ,reload, ,edit (which opens a definition in your $EDITOR) and ,help for the rest. The Clojure reader sees a comma as whitespace, so they can’t clash with real code.
  • Ctrl-C interrupts what’s running, and code that reads *in* gets its input, something REPLy never quite got right.

Without a server to connect to, neorepl starts one for the project, the way CIDER’s jack-in does, with cider-nrepl on the classpath, and stops it when you exit. It knows Leiningen, the Clojure CLI, babashka and shadow-cljs projects.

ClojureScript works too. Start a ClojureScript REPL in the session (through shadow-cljs or piggieback, which jack-in adds for you) and neorepl follows along: evaluation, completion, docs and the rest switch to ClojureScript, and :cljs/quit takes you back.

For agents

Coding agents are surprisingly good at Clojure, as long as they have a REPL to talk to. Most of them drive CLIs much better than anything else, so neorepl has a CLI that’s meant to be driven:

neorepl's CLI

  • neorepl eval evaluates code on the project’s server. Calls share a session, so *ns*, *1 and *e carry over from one call to the next, and --session keeps separate lines of work apart.
  • --format json (or edn) gives you the values, the output and the namespace as data.
  • --timeout interrupts the evaluation before giving up, so an infinite loop doesn’t take the server down with it, and --max-output keeps a stray (range) from eating the agent’s context.
  • neorepl check finds unbalanced parens, brackets and braces without a server. That’s the most common way for an agent (or a human) to break a Clojure file.
  • neorepl doc, source, apropos, complete, where, macroexpand and reload do what the REPL’s commands do.
  • neorepl server starts a server that stays up, for editors, agents and neorepl eval to share.

The exit codes are boring on purpose: 0 when all went well, 1 for an evaluation error, 3 when there’s no server to talk to and 124 for a timeout (like timeout(1)).

Playing nicely with nREPL

A new client is also a new way to break servers, or to be broken by them, and the nREPL world has quite a few servers these days. Fortunately I’ve been working on proof, a compatibility suite for nREPL, which made neorepl a good test subject.

proof checks both sides of the conversation. proof proxy sits between a client and a server and grades everything the client sends: every request has an id, need-input gets answered with stdin in the right session, and so on. proof serve is a small nREPL server that can behave like other servers in very specific ways - sending only the last value of the code like Basilisp and jank, dropping stderr, having no interrupt op, splitting output into one message per character, and so on. I ran neorepl’s REPL and CLI through both, against nREPL, babashka, Basilisp and jank, and through every quirk proof serve knows about.

neorepl came out of it in decent shape, but not unscathed:

  • Code read from stdin (neorepl eval -) that then read *in* itself crashed neorepl and left the evaluation waiting on the server for input that was never coming.
  • neorepl’s errors could leave the shell’s prompt dangling at the end of a line with servers that don’t end their error reports with a newline (hello, jank!).
  • neorepl eval exited with 1 when it lost the server, which is the status for code that failed.

All of those are fixed in 0.1. It also turned up a couple of things that are nREPL’s fault rather than neorepl’s: need-input and the end of an interrupted evaluation can go to the wrong connection when a session is used from more than one. Those will be fixed in nREPL itself.

It works the other way around too. proof now checks servers with the exact requests neorepl sends while starting its REPL and evaluating code, next to the ones CIDER, Calva, Conjure, vim-fireplace and REPLy send, so a server that breaks neorepl finds out from proof, not from a bug report. nREPL, babashka, Basilisp and jank all pass those checks today.

Making the neighbours better

Building neorepl turned up a few problems elsewhere, which is one of my favourite side effects of writing a new client. piggieback used to evaluate only the first form you sent it and silently drop the rest, and it loaded the ClojureScript compiler at startup even when nobody asked for a ClojureScript REPL, which added almost half a second to the startup of every server with ClojureScript on the classpath. The fixes for both (piggieback#162 and piggieback#161) will be in the next piggieback release. The Node.js REPL in ClojureScript itself doesn’t cope well with interrupted evaluations, and that one is next on my list.

What’s next

0.1 is the foundation. Here’s what I’d like to tackle next:

  • a test runner with proper reports
  • a full-screen inspector and stacktrace browser
  • an MCP server and a ready-made skill for coding agents
  • better Windows support

Try it

The easiest way to install neorepl is bbin:

$ bbin install io.github.nrepl/neorepl
$ cd my-project
$ neorepl

The README has the details.

It’s early days, so I’m sure you’ll find rough edges. Please file issues, and tell me what you’d like to see next. And if you’re one of the people who kept lein repl and REPLy going over the years - thank you! neorepl wouldn’t exist without you.

Keep hacking!

Permalink

Clojure 1.13.0-alpha9

Clojure 1.13.0-alpha9 is now available! Find download and usage information on the Downloads page.

  • Dropped :defaults directive - use cases better handled by others now

  • CLJ-2983 Removed :defaults tests

  • Faster transient propagation of metadata in into and update-vals

  • Revise impl and inline of meta via RT

  • CLJ-2978 Catch protocol extension of non-protocol methods

  • CLJ-2982 TransientArrayMap should expand to same threshold as PersistentArrayMap

  • Favor class files when reproducible build sets class and source timestamps the same

  • CLJ-2984 Set file timestamps to last git commit time

Try it out

Update your deps.edn :deps with:

org.clojure/clojure {:mvn/version "1.13.0-alpha9"}

Start a REPL with the Clojure CLI (any version) with:

clj -Sdeps '{:deps {org.clojure/clojure {:mvn/version "1.13.0-alpha9"}}}'

Permalink

A useful personal agent

I started building a useful personal agent. What makes it useful is that it runs on my computer, but this immediately raises concerns about data exfiltration. I use Apple reminders (among other things) to organise my life so I had it write a little script to be able to interact with those, which works great. One of the main things I want help with is getting a handle on my chaotic piles of todos. They are scattered all over the place and all mixed up. Part of the problem with my current system is that aspirational things are mixed up with hard deadlines, so on busy or unpredictable weeks things that aren’t essential just pile up and stay there and then the “today” list just keeps rolling everything over until there are dozens of things on it, which I couldn’t possibly get done in a single day if I tried. And then I’m just mentally keeping tabs on what is actually important or on a real deadline, which defeats the purpose of the system.

The problem is that since I had a baby every day is unpredictable. I used to have a pretty good handle on my life, but none of my old systems really work anymore. There’s no way to predict how my nights will go anymore so my energy levels are very inconsistent, and I can’t even know what the days will be like. Anyway sounds like I need a follow-up post on the woes of working parenthood. Point being, I need a new system and so far am having fun building an LLM-based one.

Like I said what makes the agent useful is that it runs on my computer. It has access to my data and the internet, and when it was just helping me with reminders it only had access to those, which I write, so I trust them.

The problem is that in the process of having it help me organize the reminders, I realized that the answers to most of the questions it was asking me could be found in my emails. It was still helpful and less overwhelming having a bot help me organize things, but it still needed a lot of information and context from me that it could have found in my emails if it’d had access. But the problem with giving a bot access to email is that emails come from other people. A way to inject a prompt to my bot is the missing piece of the lethal trifecta, so I need to find a way to do it safely. Turns out it’s just a hard problem, but there is some interesting research out there and I’m working on coming up with something that will work to safely allow the bot get answers from my email without a way to pass malicious or sensitive information it finds there forward to an agent that has ways to access the internet.

Permalink

Stringulation

I’ve been stringing you along for years. Time for a (breaking?) change?

Shall I compare thee to a … string?

When porting Clojure for the JVM to the CLR, the question often arises of when a particular aspect of computation is intrinsic to Clojure or just an exposure of a feature of the JVM. For the former, I try to duplicate behavior; for the latter, I try to expose the corresponding CLR behavior. I have run into this not infrequently, particularly in the early days of the porting effort.

One area where it came up very early was with string comparison. And, frankly, I did not give it a lot of thought. Where ClojureJVM used java.lang.String.compareTo, I used System.String.CompareTo. In other words, I made the decision to expose the underlying platform mechanism. In retrospect, this was probably not the right decision. These methods are significantly different.

The JVM String.compareTo is an ordinal comparison of UTF-16 strings. This is a straightforward lexicographic comparison, performed character-by-character. The CLR String.CompareTo is culture-sensitive, i.e., the result of comparing two strings varies depending on the culture in effect at that time, which is thread-dependent.

There are several consequences of using CompareTo on the CLR, of varying import.

ClojureJVM and ClojureCLR differ on things such as comparisons.

(compare "a\u00ADb" "ab") ;; => non-zero meaning not equal (JVM)
(compare "a\u00ADb" "ab") ;; => zero, meaning equal        (CLR) 

(\u00AD is the soft-hyphen character.)

And, thus, sort order:

(sort ["b" "B" "a" "A"]) ;; => ("A" "B" "a" "b")  (JVM)
(sort ["b" "B" "a" "A"]) ;; => ("a" "A" "b" "B")  (CLR)

compare and = don’t agree.

(compare "a\u00ADb" "ab") ;; => 0  (equal)     (CLR)
(= "a\u00ADb" "ab")       ;; false (not equal) (CLR)

The difference here is that compare uses String.CompareTo (culture-sensitive) and = uses String.Equals (ordinal comparison).

sorted-set and hash-set yield different sets on the same inputs.

 (count (sorted-set "a\u00ADb" "ab")) ;; => 1, one string silently dropped because it compares as equal (CLR)
 (count (hash-set "a\u00ADb" "ab"))   ;; => 2, because hash sets use ordinal = (CLR)

Culture-sensitivity bites in some related places where string comparisons occur.

For example,

(compare :b :B) ;; => 32 (positive, indicating :b > :B) (JVM)
(compare :b :B) ;; => -1 (negative, indicating :b < :B) (CLR)

Culture-sensitive string compares are thread-dependent.

For example, ASP.NET Core sets the culture per request from Accept-Language, so (sort names) silently follows each user’s collation. Picking up information from the thread is fine; I feel it is better for it to be a deliberate choice rather than delivered behind your back.

Culture-sensitive comparisons are more expensive than ordinal comparisons.

Sorting a bunch of strings or creating a sorted map under ordinal comparison executes roughly 54-61% fewer instructions (instruction counts measured on one machine, not timings) than a culture-sensitive comparison using en-US. Capturing one culture comparer at startup to avoid thread lookup of the culture decreases the instruction count by roughly 5% compared to String.CompareTo today.

Culture-sensitive comparisons have not been consistent over time.

“Before .NET 5, the .NET globalization APIs used different underlying libraries on different platforms. … If you upgrade your app to target .NET 5 or later, you might see changes in your app even if you don’t realize you’re using globalization facilities.” (See Globalization and ICU - .NET). And you can switch globalization providers. ClojureCLR on .NET Framework uses NLS, not ICU, so Framework and .NET builds of ClojureCLR on the same machine can yield different results. That’s just on Windows. Let us not discuss Linux and Mac. Read it and weep.

Going deeper

The original focus in the benchmarking investigation surfaced the compare issue outlined above. After the initial results, I decided to expand the search to all string manipulation in the ClojureCLR, with comparisons against ClojureJVM where appropriate. Most of this internal string manipulation is related to reading data (which in Lisp-land includes programs) which arguably should not be culture/locale sensitive. It is not in ClojureJVM – all is ordinal. There are a few places where culture-sensitivity snuck into the ClojureCLR code, by carelessness or ignorance. Some are so marginal that I’m guessing they have never been encountered. For example, if we let <SHY> represent the soft-hyphen character, when reading source code:

  • JVM reads foo:<SHY> => creates symbol foo:<SHY>
  • CLR reads foo:<SHY> => throws an exception under en-US (at least)

The most significant culture-dependent bugs are

  • tr-TR: the flag to turn on direct linking in the compiler is ignored – the dotless ‘i’ (U+0131) makes an appearance.
  • sv-SE: (+ 1 1E-10M) doesn’t compile. Swedish uses a different minus sign character.

I consider these examples to be bugs. Fixing them is not a breaking change.

There are five functions (starts-with?, ends-with?, index-of, last-index-of, replace-first) in the clojure.string library that are in conflict with the JVM version and also demonstrably incorrect. These bugs are visible to users today, but they are bugs and should be fixed. Example: (replace-first "a<SHY>b" "ab" "X") gives "Xb" and corrupts the string. Results would change only for strings containing ignorable characters (soft hyphen, NUL, combining marks) and the changes would match the JVM’s answers.

A proposal

I plan to make two sets of changes. The first set is to fix the bugs above: the reader, the compiler, number literals, and clojure.string. I am not <SHY> about making these changes.

The second set of changes will be more user-visible and thus have the potential to break user code. This involves changing the definition of compare to be ordinal-based. To minimize impact, we can make this change selectable at startup. Set the switch to ‘culture’ and you get the current behavior; set it to ‘ordinal’ and you get ordinal-based comparisons. The ‘ordinal’ mode is pretty much identical to ClojureJVM behavior.

There is a third possibility: A captured culture mode which at startup creates a string comparator based on the CurrentCulture in effect at system startup. That is used by the compare function. If you change CurrentCulture, it will not be seen by this; there is no thread-dependency. You don’t get the large speedup, but do get the 5% savings from the no-thread-lookup. I’m not sure it is really worth it. The audience would seem to be a culture-mode user who sorts heavily and doesn’t care about threads. I’m not sure that’s much of a market. (This would require a little more testing before full validation as an option. )

For ordinal mode users, varying culture in things such as sorting is still possible. Most of the Clojure-defined functions, such as sort and sorted-set, have variations that take a comparator function. In fact, I think the general advice should be to be explicit in your sorting and comparator options always. To make this easier, we will add a culture-comparator function that will return a comparator function based either on the value of CurrentCulture when called or a supplied culture. You could write (sorted-set-by (culture-comparator "sv-SE")).

The choices

The bug fixes will happen regardless. For the compare changes, we have some choices.

  1. Do nothing. Not really an option, from my viewpoint.
  2. Put in the switch. Default is ‘culture’. If you want performance, you have to ask for it.
  3. Put in the switch. Default is ‘ordinal’. Performance and consistency are the default.

To be clear, my own preference is #3. The current situation has internal inconsistencies, is inconsistent with the JVM, and is slower overall.

An additional choice is

  1. Do we need the ‘captured culture’ mode?

I plan to put a link to this document up on the #clr channel in the Clojurians Slack for feedback. I’ll amend this post to reflect that discussion.

The result

TBD.

AI disclaimer/acknowledgement

I’ve been using AI coding tools extensively in my benchmarking work. There is a lot of tedious coding involved and the tools do it for me a lot faster and more correctly than I would on my own. This document reports on the results of that work. I drafted the document, then used the AI coding tool to vet it for accuracy, resulting in some typo corrections and a few suggested rewordings of inaccurate phrases. A stray comment in its review of my first draft sent me down a new line of inquiry that ended up in the two-phase approach here. The instructions to the agent include not drafting any code that might go into the ClojureCLR code base. It can identify locations and make prose suggestions only. (And it can write all the testing code its non-existent heart desires. It makes measurement runs on my commits as we go along.)

Permalink

A look at the Clojure CLI REPL

Just before this year's Clojure/conj, the Clojure team released a new Clojure CLI REPL. Per the README:A REPL for the Clojure CLI featuring multi-line editing with proper indentation, bracket highlighting, structural editing, inline eval, doc lookup for Clojure and Java, a data inspector, configurable prompts, keybindings, and much more.

Permalink

Clojure in the Age of Language Models

In the age of generative models, the most important skill for a developer is to be able to recognize the shape of the problem and pick the correct way to express it. What's relevant today is the ability to do high level reasoning about algorithms, data structures, and data flows within the system. Imperative programming is quickly becoming akin to writing assembly because language models are quite competent at writing code in the small, while they stumble at high level design and architecture. So it is good to learn a language that operates at a higher level, such as Clojure, which is data-centric, composes functions declaratively, and keeps code close to the shape of the problem.

LLMs can generate code much quicker than I can, but the issue is how to test that the generated code does what I want. Before you can test anything in many languages, you have to recompile the program, and that may take quite a few minutes for larger projects. The length of that feedback cycle, in turn, sets the pace of your progress.

We also need to talk about the tedious reality of rebuilding application state. That matters even more when you are not writing the code yourself, so the output is inherently less intentional. Clojure collapses that loop because our workflow does not draw a hard line between when code is read, when it is compiled, and when it runs. Clojure runs in a live process, and you can redefine a function at the REPL and have the new version take effect immediately, without restarting. State stays in place while you change the code that operates on it. Inspecting the state to reproduce behaviors and verify fixes can be done instantly when you can reach into a running process.

An agent can similarly connect to a REPL to diagnose an issue and swap out the code without any downtime. Agents fundamentally need observability in order to get useful feedback about the changes they are making. Working with the REPL means the agent doesn’t have to go through all the steps of compiling and rebuilding the app, then logging its output to see the result. That translates into having to do fewer iterations to reach a working system, which becomes particularly valuable when working with a large codebase. The functionality of your system can keep evolving as you load new code into the running process without any restarts.

Another advantage comes from immutability, which helps reliably control the operating context of the program environment. When the majority of the logic in an application is written using pure functions, the agent can safely consider and test pieces in isolation without having to reason about the entire program. Agents can write functions one at a time, test them in the running program instead of making a whole bunch of changes, then running tests to find out if they worked.

At this point, I find anything with a compile cycle is a nonstarter if I have an option to have a live programming environment. It is not just compile and startup time that ends up being painful. A bigger problem is the tedious necessity of having to reconstruct the desired state every single run. When you have something small with limited functionality, that is fine, but as your application grows, rebuilding the state can take a significant effort. You might have to click through some menus in the user interface, wait for data to be processed from a service or a database, and so on. Being able to put your application in a particular state and make changes within that context is just a qualitatively better development experience.

Then there's the powerful macro system , allowing you to adapt the language to the problem domain, eliminating a lot of boilerplate code you would have to write otherwise. What makes this particularly powerful in Clojure is homoiconic syntax, where code is written using data structure literals. Since there is a common syntax for expressing both logic and data, a program can take any piece of code and manipulate it as it would with any other data structure, then evaluate it. This makes it incredibly easy to add new semantics because all you need is to make templates out of code. A macro works much like a function that accepts the code you wrote and produces the form that actually gets evaluated. The expression for printing a string does the printing when you evaluate it, but it is also nothing more than a list of the println symbol and the string itself.

One major advantage of S-expression based syntax is that it's easy for both humans and machines to read. And since the state can be trivially serialized as plain data structures, an agent can dump it in the REPL to inspect what’s happening in the application at any time. The data orientated nature of the language is a natural fit for LLMs since these models operate on text, making it exceptionally easy for them to see how data flows through the system without any opaque object graphs to worry about.

Furthermore, all the functions operate on a common set of data structures, allowing you to compose them together to transform data like Lego blocks. Clojure programs tend to be far more concise since the code is largely written through declarative composition of functions from the standard library, which encapsulate the implementation details. The result is that a codebase tends to be much shorter, leaving far less code to repeat. This conciseness matters since a smaller program costs fewer tokens, and fewer tokens leave more room in the context, making the language friendlier for local models.

In my experience, models lacking the broader context while making changes in a piece of code is one of the most common failure cases. Having a terse syntax means that more relevant code lives directly in the context window, directly addressing the problem. The model gets a significantly better view of what you are trying to do and can make much better decisions. If the whole call graph is sitting in the context window, the model sees how all the pieces fit together.

All these features combine to make a perfect environment for the agent to work in. Expressive syntax leads to less code repetition. Macros let you fold repeated patterns into new domain specific constructs. Code itself is just structured data that can be inspected and transformed. And the REPL ties it all together, providing you with a living system that evolves along with your code.

There are, however, a few drawbacks to the Java virtual machine, which Clojure traditionally runs on. From my experience, many developers, fairly or not, have an issue with requiring the JVM, and shy away from Clojure because of startup time, a somewhat heavyweight runtime, and perceived bootstrapping complexity.

One goal for Jolt in particular is to get more interest from outside the existing Clojure community by addressing these concerns. The compiler is a single binary, and it ships with all the tooling, such as dependency management and task running, baked in. It interops seamlessly with the native ecosystem via FFI, so you can use it in a way comparable to Python. Best of all, program distribution involves building a standalone binary similarly to Go. Thus, Jolt may remove the last big source of friction for trying Clojure.

In an age where writing code is cheap but verification and iteration remain expensive, the high level declarative style lines up exactly with what both agents and humans need to produce working code. Clojure is a great language to learn today because of the unique way it fits the era of language models.

Permalink

I wanted Jepsen's Elle in CI without a JVM, so I wrote adya

Your database documentation says REPEATABLE READ. Your code assumes it. Have you ever checked?

I spent the last stretch building adya, a black-box checker for transactional isolation. You give it a log of transactions a database ran, and it tells you which isolation guarantees held. For each one that didn't, it prints the transactions involved and the chain of reads and writes that proves the violation.

It's an independent Rust implementation of the approach from Elle (Kingsbury and Alvaro, VLDB 2020), the checker behind Jepsen's database analyses. If you know Elle, adya reads the same history formats and takes the same flags. If you don't, keep reading.

The problem with "it passed the tests"

Isolation bugs don't show up in unit tests. They need two or more transactions to interleave in one particular way, and when they happen nothing crashes. You just get a balance that's off, or two people booked into the same seat.

The classic example is write skew. Two doctors are on call, and the rule says at least one must stay on call. Each doctor's transaction checks "is the other one still on call?", sees yes, and takes themselves off. Both commit. Now nobody is on call.

Snapshot isolation allows this. Postgres's REPEATABLE READ is snapshot isolation, so it allows this. If you thought REPEATABLE READ meant "safe enough", the bug ships.

Elle's insight was that you can catch this from the outside. Record what every client asked for and what it got back, infer the dependencies between transactions, and look for cycles that a given isolation level forbids. Atul Adya's 1999 thesis catalogued those cycles (G0, G1c, G-single, G2 and so on), which is where the name comes from.

Why another implementation

Elle is excellent. It's also a Clojure library on the JVM. People who want it outside a Jepsen test end up shelling out to elle-cli from a Go or Python harness. I found two open-source projects doing exactly that in CI (barn, bytecaskdb), and the bytecaskdb PR lists what went wrong: a blocked Clojars mirror, a crash when graphviz was missing, log lines corrupting the JSON output.

On my Windows machine, elle-cli 0.1.11 hung with no output on every anomalous history I gave it. On Linux it worked, but in my CI comparison it needed more than five minutes on five of forty random 300-transaction histories, and ran out of a 6 GB heap on one more. adya checked all forty in 0.2 seconds.

Elle also leaves the workload to you. You have to write the client that generates transactions, runs them against your database and records the history. That's the part most people never get to.

So adya ships both halves in one binary:

cargo install adya

# Run a workload against Postgres at REPEATABLE READ,
# then check the history against serializability.
adya run postgres --url postgres://localhost/test -i repeatable-read -c serializable

What it looks like

You don't need a database to try it. adya has a built-in simulated database that implements isolation levels the textbook way:

$ adya run sim -i snapshot-isolation -c serializable -n 300 -p 4 --seed 1
history.jsonl   false

G2-item #0
  Let:
    T234 = {"index":234,"process":1,"type":"ok","value":[["r",10,[4]],["append",3,6],["r",3,[4,5,6]]]}
    T237 = {"index":237,"process":3,"type":"ok","value":[["append",9,13],["append",10,5],["r",3,[4,5]],["r",10,[4,5]]]}
  Then:
    - T234 < T237, because T234 did not observe T237's append of 5 to 10.
    - However, T237 < T234, because T237 did not observe T234's append of 6 to 3: a contradiction!

That's the doctors problem with list keys. T234 read key 10 before T237 appended to it, so T234 has to come first. T237 read key 3 before T234 appended to it, so T237 has to come first. Both can't be true, so no serial order exists. Snapshot isolation allows that cycle. Serializability doesn't.

How it works

The trick that makes this tractable is the workload. adya's default, borrowed from Elle, is list-append: every key holds a list, transactions append unique numbers to lists and read whole lists back. Each read then tells you the order of every append before it. If one client read [1, 2] and another read [1, 2, 5], you know 5 came after 2, without asking the database anything.

From those orders adya builds a dependency graph over transactions, using Adya's three edge types:

  • ww: T2 appended right after T1's append to the same key.
  • wr: T2 read a list ending in T1's append.
  • rw: T1 read a list that T2 later appended to, so T1 didn't see T2's write.

If you're checking a model with real-time guarantees (strict serializability), it adds edges for "T1 finished before T2 started", using a transitive reduction so the graph stays roughly linear in the history.

Then it looks for cycles, one strongly connected component at a time. Each anomaly class is a cycle with constraints. G-single has exactly one rw edge. G2-item has at least two, with two of them adjacent. G1c has no rw edges and at least one wr. Instead of enumerating cycles and classifying them afterwards, adya runs a breadth-first search over pairs of (transaction, path state). The path state is seven bits: how many rw edges so far (0, 1, or 2+), whether the last edge was rw, whether the first one was, whether two were adjacent, whether it has seen a wr, and whether it has used a real-time edge. One BFS then returns the shortest cycle of exactly the shape you asked for.

Before searching, it runs cheaper existence checks. For each model it asks whether the subgraph that model forbids cycles in has a nontrivial SCC at all. If the ww/wr-plus-one-rw subgraph is acyclic, snapshot isolation holds as far as cycles go, and adya skips every search that could only find SI violations. It also searches the most severe anomalies first and skips anything they imply. A 100,000-transaction history (30 MB of JSON) checks in about 1.6 seconds on my laptop.

Is it right?

A checker that's wrong is worse than none, so this took most of the work.

  1. Elle's own expected results. elle-cli's test suite ships 56 list-append and rw-register histories along with the JSON verdicts Elle produced for them. adya matches all 56: same verdict, same anomaly types, and the same weakest models ruled out. Getting there taught me two Elle details I would never have guessed. It labels an edge that carries several relations by a fixed priority (ww before wr before rw before real-time), but tests whether a cycle exists using any relation the edge carries. And it counts a composed edge that loops back to the same transaction as a cycle.
  2. Elle itself, live. On every push, CI generates random histories with adya's simulator and checks each with both tools. In the first full run Elle finished 34 of 40, and adya agreed on all 34.
  3. Databases with known answers. The simulator implements serializable, snapshot isolation, read committed (with write locks), read uncommitted, and a deliberately broken "snapshot" that writes back stale state. Tests assert that correct runs produce zero anomalies at their own level and the expected ones above it. That test caught a bug in my simulator, not in the checker: my first read committed took no write locks, which is weaker than any real database.

What Postgres and MySQL did

CI runs 4,000 transactions from 10 clients over 6 hot keys against Postgres 17 and MySQL 8.4, at each isolation level, and checks every history against a ladder of models. List-append results:

database, level serializable snapshot isolation read committed
Postgres READ COMMITTED G-single, G2-item, internal, lost update G-single, internal, lost update valid
Postgres REPEATABLE READ G2-item valid valid
Postgres SERIALIZABLE valid valid valid
MySQL REPEATABLE READ G-single, G2-item, internal, lost update G-single, internal, lost update valid
MySQL SERIALIZABLE valid valid valid

Postgres did what its docs say. REPEATABLE READ is snapshot isolation, write skew included, and SERIALIZABLE came out strict serializable. I also restarted Postgres every few seconds during a 6,000-transaction SERIALIZABLE run (--fault "docker restart -t 0 pg"). Nine restarts, 2,957 commits, 3,043 failures, and the history still checked clean.

MySQL's REPEATABLE READ is weaker than its name. It lets a transaction lose another's update and see part of another transaction's writes, so it isn't snapshot isolation. Jepsen reported the same thing about MySQL 8.0.34 in 2023. adya reproduces it from a cold start in a few seconds.

Testing your own database

The built-in drivers cover SQLite, Postgres and MySQL. For anything else there's adya run exec: adya starts your client once per process, writes one JSON line per transaction to its stdin, and reads back one line saying what happened.

{"value":[["append",3,7],["r",4,null]]}
{"type":"ok","value":[["append",3,7],["r",4,[1,7]]]}

The repo has a 59-line Python client for SQLite as a template. Swap the SQL and you're testing your own database.

Limits

adya covers list-append and read-write register workloads. It doesn't do predicate reads, or Elle's bank, set and counter checkers. It prints text proofs instead of Graphviz plots. A clean result means it found no anomaly in that history. It doesn't prove your database is correct, so run it longer, with fewer keys for more contention, and with faults.

If you run it against something interesting, I'd like to hear what it found. Issues and PRs are open.

Permalink

What would a useful agent be like?

I use coding agents for basically all of my day day to work now. Recently I’ve been seeing more and more consumer-facing agent-like products and have been trying them out but find none are really able to do the things I want, mostly because they are all run by sketchy tech megacorps that I don’t trust and don’t want to connect my data to. Unfortunately Siri still really sucks, but in theory it’s more in the realm of what I actually want and Apple already owns my entire digital life anyway.

It made me wonder what a useful personal agent would be like. These are at least some of the requirements:

  • It needs access to all of the things I already use. For me that means protonmail for email, apple calendar and reminders, my obsidian personal wiki, and more. I’m not migrating my digital life to accommodate a bot.
  • It needs to remember things between sessions. LLMs are stateless and that makes them bad at long running tasks without harnesses that manage context.
  • It needs to be able to carry out long running tasks, like researching things for me or doing grunge work like organizing all my stuff. Getting a handle on the total dumpster fire of notes and todos I have scattered across a dozen different tools would be legitimately useful to me.
  • It needs to be able to schedule things for itself. Not everything needs to be a reminder visible to me. Some things I want are like “check if this computer is on sale yet”, “check for discounts on flights”. There are lots of little tasks I do throughout the day that are like this but I can’t be bothered to script them.
  • It needs to live in one place but be accessible anywhere, like on my computer and phone at least. Probably I’ll want to use it from multiple computers.
  • It would be really cool if I could even share them. There are some projects like volunteer orgs I’m a part of or home renovations that I collaborate with other people on.
  • It should at least be able to collaborate with other agents.
  • It should be self-managing and self-improving. I should never have to explain something twice. Over time it would just absorb my preferences and accumulate tribal knowledge about my life and projects, like a good assistant.
  • It needs to know how to use or maybe even make apps. I hate chat as a human-computer interface, it’s too unspecific and slow for most of what I want to do with computers.

Technically I think these requirements imply at least some of:

  • It needs to be able to read and write files.
  • It needs to be installed on one computer that never sleeps and be accessible to others over my tailnet, or similar.
  • It needs access to the internet though and probably needs a browser.
  • At least some of the agents need access to my actual computer. I don’t think there’s a practical way to give them access to my Apple or protonmail accounts but they could do everything they need from my personal Mac directly.

Anyway there are probably a lot more things that will come up but thinking about this has made me realize I want to try to build this. We’ll see how it goes!

Permalink

The Pull Newsletter (Oct 5, 2026)

Welcome to the first edition of The Pull, our new monthly Datomic newsletter rounding up projects and announcements from the Datomic community and the Datomic team.

Community Projects

  • EACL v8 for Datomic Pro is now available on Clojars as RC1. EACL (Enterprise Access ControL) is a situated ReBAC authorization library inspired by SpiceDB, built in Clojure and backed by Datomic Pro, Datahike, Datalevin or DataScript, though Datomic Pro remains the primary backend. Read the announcement or try the demo.

  • Trydatomic.org, an interactive website to learn how to query a Datomic database using Datalog, just got a content and design refresh.

Datomic Team News

  • In case you missed it: In April we released Datomic 1.0.7622, a big, feature-filled changelog with several performance improvements. It includes:

    • Feature: Read-only connections to storage and backups

    • Feature: rseek-datoms, reverse index iteration, a complement to seek-datoms

    • Performance: Reduce log write amplification for systems with many transactions

    • New API: list-backups, which lists the timepoints available in a backup repository

    • Performance: Reduce CPU and memory required to calculate index metrics

    • Fix: Regression introduced in 1.0.7556 which broke the ddb-local protocol

  • Also earlier this year, Joe Lane, principal engineer at Nubank on the Datomic Core Dev team, gave a talk at the Unlocked Conference about Immutability in Motion, which explores how immutability and multi-tier caching power database performance at massive scale.

  • Recently Nubank engineers João Nascimento Mello and Mateus Oliveira, supported by engineers Gabrielle Cadurim and Carolina Silva, presented an online Day of Datomic workshop in connection with the 2026 Clojure Conj, now available to watch on YouTube.

  • We’re hosting DatomicConf this December 11, 2026 in Durham, North Carolina. Registration and CFP are now open, so join us there. Details at conf.datomic.com.

Questions or suggestions? Reach out to the Datomic team and the rest of the community at the #datomic channel on Clojurians Slack.

Permalink

Copyright © 2009, Planet Clojure. No rights reserved.
Planet Clojure is maintained by Baishamapayan Ghose.
Clojure and the Clojure logo are Copyright © 2008-2009, Rich Hickey.
Theme by Brajeshwar.