✨ $500 AI Visibility Audit — live at Spurlock Studios. Book the audit

What is filed under AI benchmarks?

Posts tagged AI benchmarks collect Will Spurlock's writing on this topic. The three newest excerpts: How I direct Kimi K2 by Moonshot AI for agentic workflows, long-context tool calling, and workflow automation. A 1 trillion parameter MoE model with competitive benchmarks at 5-17x lower cost than GPT-5 and Claude. OpenAI o3-Mini + o3-Mini-High: The First Free-Tier Reasoning ModelToday marks a genuine milestone in the democratization of AI reasoning capabilities. Ope OpenAI launches o1-preview and o1-mini — the first reasoning models trained to 'think longer' before responding. Here's what chain-of-thought AI means for builders.

5 posts as of 2026-05-01, counted from content/blog frontmatter

Tagged: AI benchmarks

Discover insights and strategies for leveraging technology in business

Frequently asked questions

What does AI benchmarks mean on this blog?

Posts tagged AI benchmarks collect Will Spurlock's writing on this topic. The three newest excerpts: How I direct Kimi K2 by Moonshot AI for agentic workflows, long-context tool calling, and workflow automation. A 1 trillion parameter MoE model with competitive benchmarks at 5-17x lower cost than GPT-5 and Claude. OpenAI o3-Mini + o3-Mini-High: The First Free-Tier Reasoning ModelToday marks a genuine milestone in the democratization of AI reasoning capabilities. Ope OpenAI launches o1-preview and o1-mini — the first reasoning models trained to 'think longer' before responding. Here's what chain-of-thought AI means for builders.

How many posts are filed under AI benchmarks?

5 posts as of 2026-05-01, counted from content/blog frontmatter.

What should I read first in AI benchmarks?

Start with "Kimi K2 Open Weights: How I Prompted Moonshot's Frontier Model for Agentic Tool Use". How I direct Kimi K2 by Moonshot AI for agentic workflows, long-context tool calling, and workflow automation. A 1 trillion parameter MoE model with competitive benchmarks at 5-17x lower cost than GPT-5 and Claude.