We gave open-source Al a real job. It had opinions

We gave open-source Al a real job. It had opinions

About the session

Open-source AI models are a little like a brilliant intern who shows up amazing on Monday and completely different on Tuesday.

This week Ahmed and Jeff put two of the most talked-about open-source models - Kimi K3 and Qwen 3.x - to work on real tasks and watched what happened. Spoiler: it got weird. In one demo, instead of just answering, the AI starts arguing with Computer about which tool to use.

Then the plot twist. The exact same AI, running on two different hosting services, went from "pretty solid" to "barely functional." Same brain, wildly different behavior - and it had nothing to do with the AI being dumb. The plumbing behind it was the problem.

And no, it's not just us - a bunch of other teams have hit the same wall.

The lesson? With open-source AI, the "does this actually work?" homework lands on you. It's not enough to pick a smart model - you have to test the whole setup it's running in.

Speakers

  • Ahmed Bashir

    Ahmed Bashir

    Chief Technology Officer, DevRev

  • JS

    Jeff Smith

    Wrangling agentic systems, benchmarking and leaderboards @ DevRev