What we measure, what we learn running agents built from a company's own knowledge, and what changes in the product.
Follow along in a feed reader: RSS feed
I tested Jev and six LLMs from OpenAI, Anthropic and DeepSeek on 30 chatbot messages. Here’s what I learned about accuracy, speed, cost and confidence.