v1.0Updated March 7, 2026Curated by Trent Jackson

AI Tools

Models, agents, and workflows. Which AI tool for what job, and why the answer keeps changing.

The Current AI Meta for Builders

The AI landscape shifts weekly. New models, new benchmarks, new claims. But for founders and builders — people shipping real products — the question isn't 'what's newest?' It's 'what actually works when the stakes are real?'

Three layers matter: the reasoning model (where you think), the coding agent (where you build), and the utility models (embeddings, image gen, classification). Most builders only need one strong pick per layer.

The Reasoning Layer

Claude owns this for builder workflows right now. The reasoning handles architectural decisions, the context window processes complex analysis, and the output doesn't drift into generic AI-speak. Pick a reasoning model you trust, then build your workflow around it.

The Coding Agent Layer

The gap between IDE copilots (suggesting lines) and agentic coding (reading your whole project, running builds, fixing errors) is the gap between autocomplete and a teammate. If you ship your own code, agentic coding is the highest-leverage AI tool available right now.

The Utility Layer

Embeddings for search, image generation for content, classification for routing. These aren't daily-driver tools — they fill specific gaps. Pick the cheapest option that meets the quality bar.

Recommendations

I Use This

Claude (Anthropic)

Free tier / $20/mo Pro / API usage-based
Tier 1

The strongest reasoning model available. Deep analysis, nuanced writing, and structured output that doesn't drift into generic patterns. The model to reach for when the thinking matters.

I've tested every major model for real work — not benchmarks, actual production tasks. Claude consistently handles the problems where reasoning quality is the bottleneck. If you're making decisions, writing strategy, or analyzing complex situations, this is the one.

Pros
  • Strongest reasoning and nuance across models
  • Consistent voice — doesn't drift into generic patterns
  • Tool use and structured output are reliable
  • Context window handles complex multi-file analysis
Cons
  • Rate limits can bottleneck batch operations
  • Opus is expensive for high-volume tasks
  • No native image generation

Best for: Strategic thinking, deep analysis, content that requires nuance

Claude Code

Included with Claude Pro / API usage-based
Tier 1

Agentic coding from the terminal. Reads your entire codebase, runs builds, fixes errors, and ships features with context that IDE copilots can't match. The difference between autocomplete and a teammate.

This changed how fast I ship. It doesn't just suggest code — it understands the full project, runs the build, and fixes its own mistakes. If you write your own code as a founder, this is the highest-leverage AI tool available right now.

Pros
  • Deep codebase awareness — reads hundreds of files
  • Runs builds and tests, fixes errors in-loop
  • Persistent memory across sessions
  • Terminal-native — no IDE dependency
Cons
  • Token-heavy for large repos
  • Learning curve for effective prompting
  • Requires trust — it can and will modify files

Best for: Founders and builders who ship their own code

Gemini (Google)

Free tier / Pay-as-you-go API
Tier 1

Google's model family. Strong for embeddings, image generation, and multimodal tasks. Not the primary reasoning model, but fills gaps at lower cost.

Gemini is a specialist pick, not a daily driver. The embeddings are fast and cheap, the image generation is solid. Use Claude for thinking, Gemini for the utility tasks around it.

Pros
  • Excellent embeddings at low cost
  • Strong image generation capabilities
  • 768-dim vectors are compact and fast
Cons
  • Reasoning doesn't match Claude for complex tasks
  • API stability has been inconsistent

Best for: Embeddings, image generation, and cost-sensitive tasks

Trusted Sources

Simon Willison

Django co-creator. Tests every model hands-on.

Ethan Mollick

Wharton professor studying AI impact with academic rigor.

Latent Space (swyx)

10M+ readers. Coined 'AI Engineer' as a role.