Ways to work together

Structured ways to get senior judgment into a decision that is already moving.

These are formats, not packages. The work is the same underneath: frame the decision, make the tradeoffs explicit, and leave the team with a path they can execute. Scope follows a discovery call.

Lowest commitment · 90 minutes · Paid

Inference working session

Use a paid 90-minute inference working session to resolve one bounded decision and identify the next test.

  • One clearly framed inference decision
  • Options and explicit trade-offs
  • The next measurement or experiment
  • A short written decision note after the call
Ask about a working session

01

Signature

1 week

Inference Readiness Review

Find out where your production inference path breaks before it breaks — quota, provider concentration, cost per request, and where the latency budget is actually spent.

Best for: AI-native teams running production inference at real volume, with a model, provider, or capacity decision arriving in the next 90 days and no written record of what was chosen or why.

  • Quota and provider risk map with named single points of failure
  • Model-substitution matrix and the eval evidence each swap needs
  • Latency budget decomposed by hop, and fallback strategy
  • 90-day sequence with named owners and decision gates

02

2-4 weeks

AI Architecture Sprint

Turn one consequential AI product or system decision into a production-ready architecture and a delivery path the team owns.

Best for: Teams with something important enough that the architecture cannot be improvised: an agent, a retrieval system, or an AI-native capability moving toward production without a senior architect in-house.

  • Target architecture and architecture decision records
  • System boundaries, guardrails, and human review paths
  • Evaluation, observability, latency, and cost plan
  • Phased delivery sequence and technical handoff

03

1-2 weeks

AI Direction Sprint

Answer which AI opportunity actually deserves to become a product initiative — and why the others can wait.

Best for: Founders and CTOs with strong market insight, several plausible AI directions, and no shared basis for choosing what deserves to ship first.

  • Opportunities scored by value, feasibility, uncertainty, and risk
  • Readiness and capability gap assessment
  • One prioritized bet with a defensible rationale
  • 90-day sequence with owners and decision gates

04

Ongoing, usually after a sprint

Fractional AI Architect

Senior architecture judgment embedded beside the founder and engineering team, before the company needs or can justify the role full time.

Best for: Startups making a steady stream of consequential technical decisions — model and vendor choices, production reviews, roadmap tradeoffs — with no one senior enough to hold the architecture.

  • Founder and architecture working sessions
  • Model, vendor, and platform decisions
  • Architecture and production reviews
  • Technical risk, evaluation standards, and team unblockers

05

3-6 weeks

AI Engineering Enablement

Turn architecture decisions into shared engineering practice, so the same lessons stop being relearned team by team.

Best for: Scale-ups already running AI in production, but with inconsistent evaluation, unclear ownership, and no shared review standard.

  • Shared architecture principles and review standards
  • Evaluation and production workflow patterns
  • Ownership model and right-sized governance
  • Office hours and adoption measures

A good fit when

  • There is real momentum and a customer or market signal
  • AI is strategically relevant, not decorative
  • Several plausible technical directions are open
  • Architecture decisions are becoming consequential
  • The team is capable but has no senior AI architecture leadership
  • The company moves too fast for a traditional consulting engagement

Probably not me

  • A chatbot or a basic automation with no product ownership behind it
  • Generic AI training disconnected from an operating change
  • Inexpensive development capacity
  • Implementation resources rather than technical judgment

One active architecture or direction sprint at a time, plus a small number of advisory relationships.