Aaron Yim · founding product manager, AI for coding at Microsoft

Decide what AI to build.
Ship it safely.

Upshift helps companies pick the AI bets worth making, then builds the evaluation and controls to ship them safely.

3M+People using agents we shipped
$200MInvoices processed by those agents
Teams trained, systems shipped

Kyndryl, Women in Innovation, Koch, Microsoft, NYU, Cornell.

Why Upshift

Most teams have no shortage of AI ideas. What they lack is a way to tell which ones are worth building, and the evidence to know whether the thing they shipped actually works. We do both: narrow the portfolio to what pays off, then build the measurement and controls that make shipping a decision rather than a gamble.

That judgment comes from shipping it. Aaron Yim was the founding product manager for Microsoft's AI for Coding — setting strategy for IntelliCode, then shipping GitHub Copilot for Visual Studio — and most recently built the finance agents at Sema4 now running inside a Fortune 50 company.

Product strategy Evaluation design Model selection Fine-tuning Red-teaming Governance Deployment controls
Portfolio triageWhich bets to fund, and which to cutWeeks 1–2
EvaluationTask suites, graders, regression gatesWeeks 2–6
ControlsGuardrails, monitoring, escalation pathsWeeks 4–8

What we do

Services

Two ways to work with us. Buy the map when you need clarity. Add embedded leadership when the change needed to happen yesterday.

Bet discovery Workflow redesign Evals Rollout & guardrails Adoption

Engagement one

AI Bet Discovery Sprint

Define the bets. Deliver the map.

For teams that need clarity before committing more budget, vendors, or headcount. We compare pilots, vendor pitches and internal ideas against revenue, margin, cycle time and risk, then cut the list to the two or three bets actually worth backing.

  • A leverage map and prioritised workflow shortlist
  • A board-ready recommendation on where to bet
  • A 90-day proof plan leadership can govern
  • The team, data and access blockers in the way

Engagement two

Embedded AI Leadership

De-risk the build. Make the change stick.

For teams who already have pilots, builders or vendors in motion and need that turned into a safe, launched change. We run the steering cadence, set the launch bar, and stay on the hook until the workflow is actually live.

  • Everything in the sprint
  • Workflow redesign and rollout steering
  • Guardrails, launch criteria and a proof model
  • Adoption support and role-based training

What that draws on

AI product strategy

Which use cases justify the investment, which are premature, and what the sequencing should be. Grounded in your data, your constraints and what the models can actually do today.

Evaluation

Task suites that reflect real usage, graders calibrated against human judgment, and regression gates in CI — so quality changes show up before your customers find them.

Training & data

Dataset design, annotation strategy and fine-tuning where a general model falls short — with the evidence to show the tuned version is genuinely better, not just different.

Controls & governance

Guardrails, red-teaming, monitoring and escalation paths sized to real risk — plus the review process and documentation that hold up to internal and external scrutiny.

Workflow redesign

Current-state map of who does what and where work waits, then a future-state design covering AI steps, human review, approvals and systems touched — with the measurement plan that shows whether it actually improved.

Adoption & enablement

Manager briefings, role-based training, launch docs and office hours — plus the adoption metrics that tell you which teams are stuck, so rollout is steered rather than hoped for.

Principles

How we operate

The commitments that shape every engagement.

Evidence over demosMeasurement first
A convincing demo is not a working product. We do not call something ready until there is a measurement that says so, and a threshold agreed before we saw the result.
We'll tell you not to build itHonest scoping
Some use cases are not ready, and some never will be. Recommending against a project is part of the work — it is cheaper than finding out two quarters in.
You keep the machineryNo lock-in
Harnesses, datasets and runbooks are handed over and documented so your team can run and extend them. The goal is your independence, not a standing invoice.
Safety sized to riskProportionate controls
Controls should match the actual consequences of failure. We right-size them so safety work speeds up shipping instead of becoming a checkbox nobody reads.
DecideA ranked portfolio with the reasoning written down.
MeasureEvaluation your team runs without us in the room.
ShipControls and launch criteria proportionate to the risk.

Who you work with

The team

Two operators, not a bench of analysts. You get the people who have shipped this before, in the room.

Aaron Yim

AI product strategy & evaluation

Founding product manager for Microsoft's AI for coding, taking the strategy from zero to 3M+ users across IntelliCode and GitHub Copilot. Defined the metrics, security testing and legal and privacy guardrails behind Microsoft's first AI agent capabilities.

  • Redesigned invoice reconciliation and procure-to-pay at Sema4.ai, shipping agents that run at 93% straight-through processing
  • Led agent discovery sprints across 10 teams in Microsoft's developer division
  • Trained 750+ developers and managers

Lillian

Enterprise workflow redesign & rollout

Brings enterprise workflow redesign, modernization and stakeholder alignment across complex transformations — the discipline that turns a good AI bet into change a large organization can actually absorb.

  • Helped brands including Nestlé and Haleon turn market and internal signal into new bets and launch decisions
  • Led workflow redesign and modernization across IBM/Kyndryl, Hertz and MetLife
  • Mainframe-to-cloud migration and messy enterprise change, made executable

FAQ

Answers

Common questions about scope, engagement shape and how we work with your team.

Strategy Evaluation Controls
What does Upshift actually do?

Two things. We help you decide which AI products are worth building, and we build the evaluation, training and controls that let you ship them safely.

Do you build the product itself?

We work alongside your engineers rather than replacing them. We build the measurement and control layer, and advise on the product — your team owns the application code.

How long is a typical engagement?

An opportunity review runs about three weeks. Evaluation build-outs are six to eight. Longer partnerships are structured as a retainer with a defined review cadence.

What do we get at the end?

Working artifacts, not a slide deck: eval harnesses in your repo, datasets, launch criteria and runbooks — documented so your team can run and extend them.

Which models do you work with?

We're model-agnostic. Selection is an output of the evaluation work, not an assumption going in, and we'll re-test when a materially better option ships.

How do you handle our data?

Under your agreements and inside your environment wherever possible. We scope data access to what the work requires and document the handling before we start.

Do you work with regulated industries?

Yes. Regulated deployments are where evaluation and controls matter most, and where the documentation trail we produce tends to pay for itself.

How do we start?

An intro call. Bring the use cases you're weighing — you'll get a straight read on which are worth pursuing whether or not we work together.

Taking on new engagements

Weighing an AI bet?
Let's pressure-test it together

Book an intro call

Prefer email? aaron@withupshift.com