Measurement (Google Analytics 4) loads only if you accept. Privacy
✨ $500 AI Visibility Audit — live at Spurlock Studios. Book the audit
What is filed under AI benchmarks?
Posts tagged AI benchmarks collect Will Spurlock's writing on this topic. The three newest excerpts: How I direct Kimi K2 by Moonshot AI for agentic workflows, long-context tool calling, and workflow automation. A 1 trillion parameter MoE model with competitive benchmarks at 5-17x lower cost than GPT-5 and Claude. OpenAI o3-Mini + o3-Mini-High: The First Free-Tier Reasoning ModelToday marks a genuine milestone in the democratization of AI reasoning capabilities. Ope OpenAI launches o1-preview and o1-mini — the first reasoning models trained to 'think longer' before responding. Here's what chain-of-thought AI means for builders.
5 posts as of 2026-05-01, counted from content/blog frontmatter
Posts tagged AI benchmarks collect Will Spurlock's writing on this topic. The three newest excerpts: How I direct Kimi K2 by Moonshot AI for agentic workflows, long-context tool calling, and workflow automation. A 1 trillion parameter MoE model with competitive benchmarks at 5-17x lower cost than GPT-5 and Claude. OpenAI o3-Mini + o3-Mini-High: The First Free-Tier Reasoning ModelToday marks a genuine milestone in the democratization of AI reasoning capabilities. Ope OpenAI launches o1-preview and o1-mini — the first reasoning models trained to 'think longer' before responding. Here's what chain-of-thought AI means for builders.
How many posts are filed under AI benchmarks?
5 posts as of 2026-05-01, counted from content/blog frontmatter.
What should I read first in AI benchmarks?
Start with "Kimi K2 Open Weights: How I Prompted Moonshot's Frontier Model for Agentic Tool Use". How I direct Kimi K2 by Moonshot AI for agentic workflows, long-context tool calling, and workflow automation. A 1 trillion parameter MoE model with competitive benchmarks at 5-17x lower cost than GPT-5 and Claude.