Home

Musings

Updates, observations, and things I'm thinking about. Platforms come and go, data stays here.

Been chewing on this one for months and it still gets me: a 100% tax can change absolutely nothing. Not a loophole, not evasion, not offshore anything. Just arithmetic.

Start here: if wealth grows by multiplying, it condenses. One holder ends up with everything. That isn't a flaw in some particular model, it's what multiplicative processes do. Which means every wealth distribution that isn't condensed is being opposed by something. Fine .. but by what? I stopped asking "which institution" and started asking about coordinates: base, rate, periodicity, threshold. And then one more that turned out to matter more than the other four put together .. realisation, the share of a period's gain your base can actually see.

Set realisation to zero and a 100% levy on flow leaves the wealth vector exactly unchanged, agent by agent. The rate never gets a turn. You can't tax what you can't see, and the rate is just a multiplier on nothing. The collapse from there is violently front-loaded too: the first five percent of the realisation axis covers 32% of the reachable range.

So the base caps the region and the rate only moves you around inside it. Everything reproduces from open code and 18 tests pin it down. Paper's here if you want it, it's 20 pages.
I pre-registered a test, ran it, and it killed my own favorite hypothesis. Then I ran it a second time to be sure it wasn't a fluke. Honestly the best thing that happened to this paper.

Here's the setup. Accounting research has spent decades ranking firms on "conservatism" .. how fast a company books bad news, and how much of that news sticks around. Two different properties: timeliness and durability. Everybody measures them. Nobody separates them, and it turns out you can't. That's the theorem. Two firms can behave completely differently .. one fast and forgetful, one slow and permanent .. and file literally identical numbers. Same series, same everything.

This is the part I want to be clear about, because it's the part people get wrong when I explain it at a bar: it is not a power problem. More data doesn't fix it. The identified set is a continuum, and every instrument the field currently uses is reading the product of two things and reporting it as one thing.

What survives is less than I wanted, but it's real, and §9 is where the registered test took my preferred reading of the null out behind the shed. Paper's here. Fair warning, it's 60 pages and it earns them.
The AutoSOTA paper is highly interesting, Li, et al decided to dispense with having a monolithic model and do a function-based approach, like how CrewAI works. I'd be interested in a counterfactual study: remove many of the guardrails and, with an intelligent enough model, provide it contours into the harness instead of hard stops. Would it revolutionize science? It may. MoE, as used here, is one approach, but I believe the team there over-constrained it because by default MoE is already constrained to a role, My thought experiment: use deterministic force functions to validate the result before accumulating it, acting as the ultimate guardrail, but leverage the non-deterministic LLM's capability to explore. Either way, it's a great paper, highly recommend reading it.
My advice on Mimi: Don't take the bait *unless* you have time to dev a proving ground. Here's why, and it has nothing to do with espionage, Wall Street or noise on reddit.com/r/vibecoding: pick your model and deduce it's SWOT against your own use case, upgrade your harnesses accordingly, and make LLMs work for you. There is so much FOMO right now about squeezing the last red penny out of inference - but from a researcher who breathes in this world 18 hours a day (not including what I dream about in my sleep), first optimize for productivity, not cost.. Cost is a second-order optimization. Maybe that's Codex. Maybe that's Kimi. Maybe that's Opus/Fable. What's good with one won't be so great with another, and we've already reached the point where Goodhart's law is in play on these models.. don't take that bait. Instead, look at your task list, try them out on a few tasks, and pick one. Marry it for awhile. Keep testing new models - and this is, if anything, strongly encouraged - but you should have a daily driver on top of your explorations.. don't blend the two or you'll spend more time in harness engineering, re-scoping and more than you will spend on getting anything done.

(Why do I say this, you ask?) The answer: LLMs, as a science and art, is very young and has a long way to go. You'll be ping-ponging for quite awhile if you narrow it down to cost, since Sonnet 7 will beat Fable 5 soon at a $20 tier; that's not an idle prediction, that's a fact. In the end, the "super-intelligence" will be a commodity, we all know this in AI research. That's not what you need. You just need to get your work done. Don't be so damn cheap (because trust me, you'll spend more of your clock-time instead, and the clock doesn't give refunds). Instead, prioritize your to-do list over the noise.

Magic Recipe: 95% of your LLM activity = getting your work & life done with your favorite model (and it could be your favorite for a variety of reasons you can't explain, this is why you are human; your gut knows more than you do). Take the other 5% of your day or week and poke around at new models. If you want to go full monty with the bleeding edge, make this percentage 75% / 25%, where 75% of your life is spent with your daily driver LLM and 25% is poking at what's new. But there's a reason why it's called bleeding edge: don't kill your productivity with 1,000 papercuts.
Today I screwed up on something I should have known better on. Like the rest of the many early adopters, I have automated agentic workflows plumbed all over.. N8N, launchd, etc. "Let's save some API money," I said to myself. So off with my deepinfra.com api key with an account loaded with $50, I tried both open weight LLMs and SLMs to do what an LLM did daily in my daily comms summarizing workflow. HALLUCINATION NATION. Though points for Gemma - the latest 31B is actually close to LLM-level on a long text summarization task. But there's so much hype around open models, even the people who work in AI research, myself included, get FOMO that maybe we can save some pennies on our household automations. Nope. You still get what you pay for: SLMs and open weight LLM models by non-frontier labs are *not* production ready (yet). They are at least 18-24 months behind the leaders..
Fable? Dario's team is *terrible* at branding but *world class* when it comes to the latest LLMs. Redesigned an entire doc/github repo with 2-3 prompts this AM on Fable + Claude desktop. Back in 2023 (way back then) this would have been a 2-week project done by hand. We're living in the future! Despite the nomenclature used for model names sounding like they've arrived out of an English Lit textbook from 1987.
(NVDA) NVIDIA's biggest news on their investor call? CPUs. No, not the Nano line. Think like Apple silicon. Faint analyst coverage. Yet this is bigger news than Chinese exports or even Vera Rubin for them. $200B-$1T+ TAM for their taking. Especially if they move to RISC-V in the future.
HiPEAC has a call for posters; see you in London! https://www.hipeac.net/~jasoncbraatz/#/

A vision-language model fine-tuned on your CAD library and a year of factory-floor footage is a digital-twin starter kit on rails. It catches the gap between as-designed and as-built — the gap that eats manufacturing margin — without instrumenting a single new sensor.

Kolmogorov–Arnold networks run slower than MLPs at inference. In regulated industries — banking, pharma, defense — that trade is worth it: KAN edges are fixed, learnable functions you can read. MLP weights are an opaque chord. When auditors come knocking, an open book beats a black box.

A quantized 12B-parameter SLM with full PubMed RAG runs on a Jetson-class board. Doctors Without Borders, off-grid, 400:1 patient ratios, no internet. The frontier is not always where you think it is — sometimes it is exactly where the cloud cannot reach.

Most SMB AI projects start life as fine-tuning and end as RAG. Fine-tuning teaches the model to sound like you; RAG teaches it to know what you know. Almost everyone needs the second. They rarely need the first.

An agent that calls five tools to do what a single SQL query could do is not intelligence — it is overhead. The most useful agentic systems I have shipped are 80% deterministic plumbing and 20% LLM, in that order.

Differentiable physics lets you backprop through a manufacturing line. The same gradient that trains a neural net can now tune a production schedule. Operations research, but with a calculus you can actually ship.

Hot take: monospace fonts are underrated for professional sites. They signal I build things without saying a word.

Today

Just shipped my new personal site. Built with Next.js, Tiptap, and Postgres — deployed on the same Linode that runs my n8n instance. Caddy handles the multi-site routing. Trust through minimalism.
Christmas is almost here..