The Gospel According to Graybeard

Model flipping

My use of AI has been primarily through APIs. This is why I love Cline. It's a perfect fit (in my opinion) for human-in-the-loop, model-agnostic work (though, a recent release put its most valuable feature—diff-level edits—behind a feature flag which has me thinking about just DIY'ing my own harness).

Early on, most of my usage was fixed to a singular model for anywhere from a few weeks to months until I was compelled to "upgrade" to a newer model due to improved performance claims. Lately, though, I've gotten a lot better abut model flipping.

A good example of how I approached this is on the site/backend I just finished up for a cleaning business my wife and I recently started.

For most of the tasks that firmly fell into the "yeah, I don't want to do that by hand" column, I was leveraging Kimi K3. But for some tasks, no matter what tricks I tried (i.e., skills tweaks, repetition of goals, etc), it would just crash out and get stuck. In particular in relation to tool use and properly calling Cline's new_task tool (I'd say it properly calls it roughly ~60% of the time).

So, instead of fighting, I just flipped the model to Fable (5, 1m).

Whether the issue was tool use, getting a plan right, or writing an implementation, I found that despite costing more, the gap in model quality (anecdotally, Fable outperforming K3 on a lot of granular, "big think" style tasks) was big enough to warrant spending a bit more on specific tasks.

This morning, I had to fix a UI bug involving form validation. I went with my K3 default, tried to figure out the issue...no go. A quick flip to Fable had it resolved in a minute or two.

The point: a good skill to develop around LLMs is knowing when it's time to kick over to another model—even if briefly—to just get a task complete. This is mostly a "duh," but I've found it does take some discipline and judgment to know precisely when a task is a "flipper," versus something perfectly within the wheelhouse of a cheaper model.

Making a note here because it can be a bit maddening to know how best to use this stuff. For what it's worth, being unceremonious about what model, what lab, etc., has produced far better results than when I treated the problem like a "me" problem versus model limits.

YMMV.