My advice on Mimi: Don't take the bait *unless* you have time to dev a proving ground. Here's why, and it has nothing to do with espionage, Wall Street or noise on
reddit.com/r/vibecoding: pick your model and deduce it's SWOT against your own use case, upgrade your harnesses accordingly, and make LLMs work for you. There is so much FOMO right now about squeezing the last red penny out of inference - but from a researcher who breathes in this world 18 hours a day (not including what I dream about in my sleep),
first optimize for productivity, not cost.. Cost is a second-order optimization. Maybe that's Codex. Maybe that's Kimi. Maybe that's Opus/Fable. What's good with one won't be so great with another, and we've already reached the point where Goodhart's law is in play on these models.. don't take that bait. Instead, look at your task list, try them out on a few tasks, and pick one. Marry it for awhile. Keep testing new models - and this is, if anything, strongly encouraged - but you should have a
daily driver on top of your
explorations.. don't blend the two or you'll spend more time in harness engineering, re-scoping and more than you will spend on getting anything done.
(Why do I say this, you ask?) The answer: LLMs, as a science and art, is
very young and has a long way to go. You'll be ping-ponging for quite awhile if you narrow it down to cost, since Sonnet 7 will beat Fable 5 soon at a $20 tier; that's not an idle prediction, that's a fact. In the end, the "super-intelligence" will be a commodity, we all know this in AI research. That's not what you need. You just need to get your work done. Don't be so damn cheap (because trust me, you'll spend more of your clock-time instead, and the clock doesn't give refunds). Instead, prioritize your to-do list over the noise.
Magic Recipe: 95% of your LLM activity = getting your work & life done with your favorite model (and it could be your favorite for a variety of reasons you can't explain, this is why
you are human; your gut knows more than you do). Take the other 5% of your day or week and poke around at new models. If you want to go full monty with the bleeding edge, make this percentage 75% / 25%, where 75% of your life is spent with your daily driver LLM and 25% is poking at what's new. But there's a reason why it's called
bleeding edge: don't kill your productivity with 1,000 papercuts.