Skip to content
Work with me · Remote · Los Angeles

Ship the smallest reliable version first.

I help teams turn vague AI ideas into working systems with clear scope, measured quality, and a delivery a team can keep running. The goal is not a demo. The goal is something you can trust, evaluate, and keep improving.

What it costs

Engagements start around $5,000 for a scoped first sprint: enough to ship one end-to-end slice with real data and a measurement you can argue with. Longer builds and advisory retainers are quoted once the scope is real.

I don't quote a number before understanding the constraint, and I'll say so if the smallest useful version costs less than you expected.

Best fit
  • You need grounded answers, not a flashy demo.
  • You care about reliability, cost, and failure modes.
  • You want clear scope and fast iteration with real checkpoints.
First sprint
  1. Align on the goal, the user, and the constraint that matters most.
  2. Ship the smallest end-to-end workflow with real data and measurable output.
  3. Add guardrails, instrumentation, and a rollout path that a team can keep running.
What I usually build
  • RAG and search systems with citations, evals, and observability.
  • Agent workflows and automation that reduce manual operational work.
  • Architecture reviews for LLM apps that need lower latency and better reliability.
Where projects break
  • The prototype works, but nobody defined the real acceptance criteria.
  • Costs rise because retrieval, prompts, and failure handling stayed ad hoc.
  • The team has no clean path from pilot to production ownership.
What you leave with
  • A working slice of the system that can be shown, tested, and measured.
  • A clearer roadmap for what to automate, what to delay, and what to monitor.
  • Documentation and decisions that make the next handoff cleaner.
Availability

Taking on new consulting and build engagements. Remote-first, with short advisory work and hands-on implementation both available.

Process

How I work

Not a methodology. Five things I actually do, each one because skipping it cost me something.

Principles
  • A change that has not been run is a guess.
  • State the bet so it can be proven wrong.
  • Narrow beats clever when it runs unattended.
  1. 01

    Write the bet down before the code

    Verso exists because logging a gallery visit gives you ~15 events a year and logging each artwork gives you 150+. That claim went in the PRD first, so it could be proven wrong rather than quietly assumed.

  2. 02

    Give every change a way to prove itself

    A change that has not been run is a guess. Boot the server, render the page, run the script on real input — the failure mode of confident output is that nobody executed it.

  3. 03

    Compute the decision, never eyeball it

    merge-gate classifies a pull request from the shape of its diff. When judgement was done by eye it armed a 197-file change to land unreviewed. Automation earns trust by being narrower than a human, not broader.

  4. 04

    Fail loudly, not silently

    Cocoon warns when a site's layout changes and its rules stop matching. For an accessibility tool a silent no-op is the worst possible outcome — the user assumes it is working and it is not.

  5. 05

    Bound the blast radius

    Cocoon scopes host permissions to exactly seven domains instead of <all_urls>. FraudStream masks the card number before anything is written. Decide what the system may touch before deciding what it does.

Services

What the work looks like

Scope, approach, and what I actually ship, for each kind of engagement.

01

AI agents and automation

Agents built around your task, not a template

What I build

I build agents that do a named job: read a pile of documents and answer from them with the source attached, or take a repetitive step out of a workflow. Each one is scoped to your task, not assembled from a template.

How it works

We start with a conversation about what you're trying to accomplish. Then I prototype quickly, iterate based on your feedback, and deliver something production-ready. No 50-page proposals - just working software.

Use cases

The work I have shipped is document question-answering over a client's own corpus, fine-tuned classification, and retrieval that cites its sources. If your repetitive task is not one of those, say so on the call and I will tell you whether it is a fit.

Getting started

Book a free 30-minute call and tell me what you're working on. I'll give you honest feedback on whether AI is the right solution and what it would take to build. No sales pitch, just straight talk.

02

RAG and search systems

Search that answers from your own documents

What I build

Retrieval systems that let you ask questions of your own documents and knowledge bases. Every answer cites the passage it came from, so a wrong one is traceable rather than mysterious.

Tech stack

What I have run in production: Neon pgvector for retrieval, OpenAI for embeddings and generation, Vertex AI and BigQuery for inference at volume. If your stack is different I will learn it, but I will not pretend I have already shipped on it.

Common projects

Internal knowledge search, documentation search that answers in prose, and research corpora you need to ask questions of. The chat on this site is the public example: it answers from every essay I have published and shows you what it read.

What you get

Retrieval that answers from your own documents with the source attached. I do not publish client outcome percentages: I have not run the controlled comparison that would make such a number mean anything.

03

Technical consulting

Get unstuck on AI/ML projects with hands-on help

How I help

Sometimes you don't need someone to build the whole thing - you just need expertise to unblock your team. I do code reviews, architecture sessions, pair programming, and strategic planning for AI projects.

Common asks

"Our RAG system is returning garbage." "We need to add AI features but don't know where to start." "Our LLM costs are out of control." "Should we fine-tune or use prompting?" These are the questions the work usually starts from.

Engagement types

One-time deep dives, weekly office hours, or embedded support with your team. Flexible arrangements based on what you actually need. Remote-friendly, async-friendly.

Background

Production retrieval and classification systems at Sizzle and for Upwork clients, a plant-microbe prediction model at JGI, and the search that runs this site. The work history on the experience page lists each with its numbers.

Questions people ask first

Delivery, scope, and how we would talk to each other.

Still have questions?

Book a short call and I can tell you exactly what I would do first.

Schedule a call →
Prefer async?

Send the goal, the user, the data sources, and the constraint that matters most. I'll tell you what I would de-risk first and whether the scope makes sense. Want the background first? My skills and work history live on the experience page.