AI Radar
Simon Willisonmoonshotai/Kimi-K3Moonshot adds a commercial attribution clause to the MIT license for Kimi K3, restricting large-scale commercial use.Simon WillisonAn opinionated guide to which AI to use to do stuffEthan Mollick shifts his guide from chat models to agentic systems, dropping Gemini due to Google's lack of agentic capability.arXiv cs.AIFlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable SkillsFlowEvo compiles transient inference-time workflows into reusable executable skills without training, addressing the lack of memory in LLM agents.arXiv cs.AIRisk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk SignalsStandard ML metrics fail for wildfire systems; a monotonic framework is needed to ensure risk scores correlate with operational load.arXiv cs.AISecuring Multimodal AI through Internal Information DecompositionMultimodal attacks evade unimodal safeguards, requiring detection via cross-modal consistency rather than inspecting individual input modalities.Simon WillisonAn Inside Look at the Relay Market Powering Token Resellers and FraudA relay market in China pools abused API keys and stolen credentials to resell discounted LLM tokens, creating fraud vectors.Simon Willisonmoonshotai/Kimi-K3Moonshot adds a commercial attribution clause to the MIT license for Kimi K3, restricting large-scale commercial use.Simon WillisonAn opinionated guide to which AI to use to do stuffEthan Mollick shifts his guide from chat models to agentic systems, dropping Gemini due to Google's lack of agentic capability.arXiv cs.AIFlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable SkillsFlowEvo compiles transient inference-time workflows into reusable executable skills without training, addressing the lack of memory in LLM agents.arXiv cs.AIRisk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk SignalsStandard ML metrics fail for wildfire systems; a monotonic framework is needed to ensure risk scores correlate with operational load.arXiv cs.AISecuring Multimodal AI through Internal Information DecompositionMultimodal attacks evade unimodal safeguards, requiring detection via cross-modal consistency rather than inspecting individual input modalities.Simon WillisonAn Inside Look at the Relay Market Powering Token Resellers and FraudA relay market in China pools abused API keys and stolen credentials to resell discounted LLM tokens, creating fraud vectors.
Agentic AI · Systems Leadership

Harness engineering
for agentic AI —
at scale.

Long Yang leads Intuit's agentic AI transformation and has independently architected four solo-built multi-agent systems in production — trading, market intelligence, video, and personal finance — plus the protocol that connects them, each engineered for reliability under non-determinism, not just a demo.

01 / About

Built for reliability.

Anyone can wire up an agent that demos well — the hard part is everything around the model. Long treats evals as the ship gate, context as a first-class engineering concern, and cost the same way he engineers for correctness.

Agentic AI
Harness engineering

One of the earliest engineers to put AI agents into production at Intuit — now leading that work end to end: MCP, multi-agent architectures, and looping engineering, where agents close their own feedback loop instead of being re-prompted by hand.

Systems
A decade in production

Kafka, Snowflake, FastAPI, pgvector — a decade of owning distributed systems end-to-end before "AI engineer" was a title. Agentic AI is that same discipline applied to a harder problem: reliability under non-determinism, at scale.

Leadership
Team leadership · org-wide

Leads and mentors an engineering team at Intuit. Sets technical direction for orchestration, observability, and crash-recovery patterns — adopted across the entire GTMT AI roadmap, not just his own team's work.

02 / Projects

Four systems, one person.

Solo-architected, solo-built, running in production — evenings and weekends, alongside the full-time role above. Each one engineered to survive a bad decision, not just make good ones.

Trader Joe
Autonomous Trading

A five-role agent loop — analyst, trader, challenger, evaluator, reviewer — trading real capital on live market data.

The hard part: hard risk limits enforced in code the model can't override, and a self-improving loop where a repeated mistake gets promoted from a logged lesson into a permanent, hard-coded gate.

Read the case study
Family CFO
Autonomous Finance

A financial agent that gives a household real planning power over its own money — budgets, net worth, cash-flow forecasts, anomaly detection — running unattended against live data from any Plaid-supported bank.

The hard part: turning messy real-world transaction data into numbers a family can trust every day, with no one checking its work.

Read the case study
TMZ
Market Intelligence

A ReAct orchestrator scheduling five replaceable market-intelligence experts — collection, enrichment, novelty, evaluation, signal — over a production pipeline Trader Joe already reads from today.

The hard part: re-architecting a live production pipeline in place — every expert ships first as a byte-identical wrapper, proven against a consistency test, before any orchestration logic is allowed to change what actually runs.

Read the case study
Video Factory
Multi-Agent Production

40 modules, one GPU orchestrator, two production modes: fully autonomous for trending news — and for the deeply personal work, an AI that interviews a family, live, by voice, to fill in the gaps in their own story.

The hard part: an interview agent that has to ask a family the exact right question — grounded in real gaps in their own photo timeline, never an invented memory — plus a custom LLM router and MCP tool servers built from scratch to run the other 40 modules underneath it.

Read the case study
02b / Ecosystem

How they connect

Four independent systems, four separate reasons to exist. Wall Street is the protocol that lets them optionally exchange data and consult each other — a real contract between independently-owned systems, not just a hub node in a diagram, where every connection is opt-in and every system keeps working exactly as designed if its peers go dark.

Wall Street3 agentsFamily CFO6 agentsTraderJoe5 agentsTMZ5 agentsVideo Factory26+ agentsStandalone by design
How these systems collaborate
  • Mutual support, never mutual dependency: every system's core function works with all other peers offline.
  • Any peer being unreachable degrades to an explicit 'not connected' state — it never crashes or blocks the caller.
  • Cross-system calls are read-only by default; remote conclusions can only produce proposed actions gated by local user confirmation.
  • No shared databases across systems — HTTP contracts only, and the contract schema is the privacy whitelist.
  • Every capability is env-gated and off by default; off means byte-identical behavior to before integration.
Read the full design
03 / Experience

A decade in production.

Built

Architected Intuit's agentic platform from scratch — multi-agent LangGraph orchestration, an in-house Agent Service, an MCP server exposing data pipelines as first-class LLM tools. The E2E testing system built on top of it replaced days of manual QA with minutes of autonomous execution.

Agentic AI Engineer

IntuitNov 2025 – Present

Leading the transformation of Intuit's Go-to-Market data platform from traditional ETL to an AI-native, agentic architecture. Architected the team's agentic platform — multi-agent LangGraph orchestration, an in-house Agent Service, and an MCP server exposing data pipelines as first-class LLM tools. Built a multi-agent E2E testing system that replaced days of manual QA with minutes of autonomous execution.

Senior Software Engineer

Capital OneJul 2023 – Jun 2025

Migrated legacy infrastructure to AWS CDK and CloudFormation. Built an AI-powered CI/CD pipeline using retrieval-augmented generation to cut pipeline-failure resolution time.

Senior Software Engineer

SVB Financial GroupJun 2020 – Jun 2023

Led design of mission-critical information-reporting systems for next-generation online banking. Built and managed a team of 7+ developers.

Senior Software Engineer

CiscoJan 2015 – Jun 2020

Designed Chassis-View, an industry-first graphics-oriented network device management system. Led the migration from a monolithic architecture to microservices.

Long has been engineering and shipping software in Silicon Valley since 2015 — infrastructure, data platforms, and now agentic AI. Outside of Intuit, the four systems above, and the protocol that connects them, he's quietly incubating a separate consumer venture, Game On — a different kind of bet, on its own clock.

He believes AI, built thoughtfully, will make the world meaningfully better — more productive, more creative, more human.

— Résumé & Contact

Let's talk.

Recruiter, investor, or curious about one of the systems above — the fastest way in is the résumé or a direct email. The chat assistant in the corner can also answer most questions directly.

LinkedIn ↗[email protected]