AI reliability experiments
These are a couple of small evaluations of open-source research projects from Stanford: OpenJarvis and CollabSkill. I didn’t build either one, and nothing here was contributed back to them. I chose the questions and reviewed the results; the experiments themselves were run with coding agents. What interested me was the gap between a system’s headline number and how it actually behaves.
A better transcriber, more wrong actions
OpenJarvis runs personal AI assistants on your own device. The experiment swapped its speech recognizer for a newer one and kept everything after transcription the same. Across 28 synthetic voice clips, the newer recognizer made fewer transcription mistakes (word error rate fell from about 21% to 17%), but the assistant ended up doing something other than what was asked in 7 clips instead of 4.
One example: “check ticket OPS four two” came back as “ops for two.” That’s closer to what was said, and it pointed confidently at the wrong ticket, where the older recognizer’s garbled version had simply made the assistant ask for clarification. The voices were synthetic and the actions simulated, so this isn’t a claim about real-world harm. It’s a small case where a better component score made the overall behavior worse.
How settled is a leaderboard?
CollabSkill rates AI agents by how well they work with people on shared tasks. The reproduction recomputed its published ratings under 160 reasonable variations of the modeling choices. The top agent stayed first in all 160. The full five-agent order held in 123, because the agents in the middle are close enough that small choices swap them.