Two curves moved this year: how long an AI agent can work without supervision, and how many robots are doing paid work.
Both are further along than the sceptics allow, and further behind than the pitch decks suggest. Here’s where we think the line actually sits.
1. Autonomy is compounding faster than adoption
The most useful number in AI right now isn’t a benchmark score. It’s duration, meaning how long a task can run before an agent loses the thread.
METR has tracked this since 2019 by timing human experts on the same tasks it gives to models. Their finding: the length of task a frontier agent completes with 50% reliability doubles roughly every seven months. Since late 2023, that doubling has compressed to about four and a half months. By February 2026, the leading model cleared roughly 14.5 hours, a coin-flip chance of finishing work that would take a skilled engineer most of a day. Eighteen months earlier, the same number was in minutes.
Two caveats matter more than the headline:
50% isn’t a shipping standard. The gap between coin-flip and dependable is where most enterprise disappointment lives.
The tasks are clean. Real work arrives with prior context and half-documented systems. Performance degrades sharply as history accumulates. Agents that succeed 40 to 50% of the time on a short task can fall below 10% once the same task sits inside a long interaction.
2. Where agents earn their keep
Adoption surveys read like two different industries. PwC found roughly four in five organisations using agents in some form. Other analyses put genuine production deployment closer to one in ten. Both are true. The gap is pilots that never crossed into a system of record.
Working now:
Customer support triage. Fastest ROI, because ticket volume is large and resolution rate is trivially measurable. Salesforce reports customers automating around 70% of tier-one queries end to end.
Software engineering. The one category where the time-horizon curve translates directly into billable output.
Back-office finance. Invoicing, reconciliation, compliance monitoring. High volume, high error cost, data already structured.
Agentic checkout. ACP, AP2, UCP and x402 have split the problem into checkout, consent and settlement. Visa and Mastercard both shipped protocol-agnostic on-ramps this year.
Not working yet:
Long-horizon autonomy in messy systems. Around 70% of developers report integration problems, and most teams discover their data foundations are inadequate only after launch.
Unscoped multi-agent orchestration. Gartner expects more than 40% of agentic projects to be cancelled by end-2027 on cost, unclear value or weak risk controls.
Agent lifecycle management. Few organisations can say how many agents they’re running, who owns them, or when to retire one. Governance is the constraint on scale, not capability.
Ambiguous-ROI functions. Strategy and creative workflows demo beautifully and defend budget poorly. They shouldn’t be the first deployment.
Our read: the winning pattern in 2026 is a narrow agent, wired into one system of record, measured on one number that already appears in a board pack. Breadth is what kills these programmes.
This is why our agent exposure sits where the number is already being counted. Yenmo (YC W24) is the clearest version. India has around 65 million investors who, when they need cash, either liquidate holdings and lose the compounding plus up to 35% in capital gains tax, or take a personal loan at over 18%. More than 30% of personal loan borrowers there already hold investments and are paying roughly double the rate they’d qualify for. Yenmo lets them pledge those holdings digitally and draw at a fixed 10.5% in under five minutes, and runs agents internally across support, ops and distribution rather than selling autonomy as a product. Portfolio Pilot works the same way in personal finance: a registered investment adviser rather than a chatbot, which forces the human-in-the-loop question to be answered on day one instead of after the first complaint. Acceler8 sits in the front office, where pipeline is the number the customer already reports. None of them are trying to orchestrate a company. Each is trying to own one workflow completely.
3. Robotics stopped being a demo reel
The software unlock is the vision-language-action model, a single stack joining vision, language and motor control, so a robot can follow an instruction it was never explicitly programmed for.
Physical Intelligence’s π-series generalises across hardware platforms. NVIDIA’s GR00T pairs a reasoning backbone with a diffusion action module and reports about 40% better task success when synthetic motion data supplements real demonstrations. Google DeepMind’s Gemini Robotics added on-device inference, and latency and connectivity are deployment blockers rather than research ones. Figure’s Helix now drives full-body control in unfamiliar environments.
The commercial numbers followed:
Humanoid shipments went from about 3,000 units in 2024 to about 13,000 in 2025, with 50,000+ expected in commercial operation in 2026.
Twelve platforms are now purchasable or leasable, against three in 2024, priced from $28,000 for torso-only systems up to $245,000 for full bipeds.
Humanoid startups alone raised $8.6bn this year.
What’s genuinely proven stays narrow: defined tasks, controlled settings, vendor support on site. Agility’s Digit has moved over 100,000 totes in a live GXO deployment, which is accumulated-cycle evidence worth more than any keynote. Xiaomi has humanoids loading fasteners on its own vehicle line. Homes remain five-plus years away, and dexterous manipulation is still hardware-bound.
Cheap hardware is evidence of access, not of labour substitution.
The investable layer isn’t the robot. It’s proprietary deployment data, collected on the customer’s hardware, in the customer’s environment, on the customer’s task distribution. Open-X now holds over a million annotated demonstrations, which makes the open baseline a commodity and the private dataset the moat.
Allus AI is our version of that bet. Rather than selling a model into manufacturing, it puts an edge agent on the factory floor, which means every install adds task-specific data that no competitor can buy from a public dataset. The hardware is not the product. It is the collection mechanism.
4. What this actually looks like on the ground
Strip away the model releases and the funding rounds, and a simpler test remains. Does the thing remove a real burden from a real person, and can somebody count it?
A patient who cannot describe their pain. Around 26 million people in the US have limited English proficiency. When something goes wrong during their care, they are meaningfully more likely to be physically harmed than an English-speaking patient, 49.1% against 29.5% in the underlying study. Human interpreters solve this and cost roughly $2.95 a minute, which is why most clinics ration them. Our portfolio company Opalite Health runs real-time AI interpretation across 150+ languages instead, validated head to head against certified medical interpreters in a blinded evaluation with physician-researchers from Johns Hopkins and the VA. At CIFC Health, the reported result was 22% less time per interpreted visit at half the cost. That is not a productivity metric. It is a patient being understood on the first visit rather than the third.
A procedure that fails half the time on the first try. Central line placement is guided by 2D ultrasound, which means the clinician moves a probe while mentally reconstructing three-dimensional anatomy from flat slices. Roughly half of first attempts fail, especially among newer clinicians, and hospitals spend millions annually on training. Lumius builds real-time 3D ultrasound at a fraction of the cost of existing 3D systems. The problem being solved is a failed needle stick on a person who is already unwell.
Work that injures the people who do it. Agility’s Digit has moved more than 100,000 totes in a live GXO deployment. Repetitive lifting is where warehouse injuries come from, and it is also the task humans are worst at sustaining across a shift. Early humanoid deployments report payback in 18 to 36 months, with the economics improving sharply in high labour cost regions and for hazardous or badly understaffed work. That is the honest shape of the opportunity: not replacement, but the jobs nobody can staff.
Care that has nowhere near enough people. The WHO projects a global shortfall of 10 million health workers by 2030. Japan is running the sharpest version of this experiment, treating automation as economic necessity rather than efficiency gain. Robots are not going to close that gap, but hospital logistics, medication delivery and monitoring are exactly the kind of bounded tasks that free a nurse for the part of the job only a nurse can do.
A diagnosis that arrives too late to change anything. Cancer treatment decisions still hinge on tissue analysis that is slow, subjective between pathologists, and unevenly available depending on where a patient lives. Digistain applies AI to cancer diagnostics and is working through clinical validation and health-system entry across several markets rather than optimising a benchmark. Anto Biosciences sits further upstream, using AI on microbiome biology to find drug targets, with published work and pharma pilots as the proof points. Both are answering the same question from different ends: why does a person wait months for an answer that biology already contains?
Money that costs the most for the people with the least. Cross-border remittance and everyday banking remain expensive and slow precisely for migrant and diaspora communities, who can least absorb the spread. SpotPay send, spend money everywhere through stablecoin platform. Yenmo attacks the adjacent problem in India, where investors needing short-term cash face two bad options: sell the holdings and forfeit both the compounding and up to 35% in capital gains tax, or borrow unsecured at over 18%. Neither company is an AI story first. Both are distribution problems where automation lowers the cost of serving small accounts enough to make the unit economics work at all.
Professional work that nobody wants to do. BakerHostetler reports a 60% reduction in legal research hours using agents. BDO Colombia reports roughly half the workload across several administrative workflows. These are unglamorous numbers, and they are the reason agents keep their budgets when the pilots elsewhere get cancelled. Paratus Health is building in the same register, attacking the go-to-market data problem that makes selling into healthcare slower and more expensive than it needs to be.
The pattern across all of these: a bounded task, a person who was carrying it, and a number somebody can put in front of a board.
5. Three constraints that decide the next 36 months
Power, not silicon. Grid connection has overtaken GPU supply as the limiting factor. New large-load interconnection runs 24 to 72 months, and five to seven years in constrained regions. Roughly 2,300 GW sits in US queues, and Dominion alone has around 70 GW of large-load requests against a 24 GW all-time peak. Behind-the-meter generation isn’t a hedge anymore. It’s the schedule. Voxel Energy is our position here, building compute behind solar and storage using repurposed EV battery packs, which solves two problems at once: it sidesteps the queue, and it gives a second life to packs that would otherwise be a waste stream.
Data, not algorithms. In drug discovery and industrial AI alike, the blocker is curated, permissioned data, not model architecture. Around two-thirds of executives name data quality and governance as the reason AI initiatives fail. This is the reason we favour companies that generate their own: Allus AI on the factory floor, Opalite across interpreted clinical sessions, Anto through wet-lab work rather than public corpora.
Proof, not promise. AI drug discovery has moved from benchmark scores to wet-lab validation. One generative programme nominated a preclinical candidate after screening 78 molecules rather than thousands, in 18 months and at under a tenth of typical cost. GSK committed $50m upfront to NOETIK, and Lilly took access to Chai’s design models. Against that, Recursion discontinued REC-994 after Phase II failed to confirm earlier signals. A first approval of an AI-designed therapeutic looks most likely in 2027 or 2028, and even then it validates the tool rather than any systematic improvement in clinical success rates. We hold Digistain and Anto against exactly this standard, which is why we track their clinical readouts and peer-reviewed publications rather than their model metrics.
6. How our portfolio sits on the same map
We build across uncorrelated sectors and geographies deliberately. Each line below is the problem first, the company second, because that’s the order we underwrite in.
Cancer diagnosis is slow, subjective and unevenly available. Digistain applies AI to cancer diagnostics, moving through clinical validation and health-system entry across several markets. (MedTech, UK)
Drug targets hide in biology we can’t read at scale. Anto Biosciences does AI drug discovery on microbiome data, with a peer-reviewed publication track and pharma pilots. (Biotech, AI)
Acceler8 uses agentic AI to make workforce planning continuous — helping companies understand which roles to automate, where to redeploy talent, and who should step into critical roles. Selected for a16z Speedrun. (Agentic AI / Workforce Intelligence)
Patients who don’t speak English get worse care and clinics ration interpreters. Opalite Health (YC-backed) runs real-time AI medical interpretation across 150+ languages, HIPAA compliant and SOC 2 Type II, with sentence-level safety checks and notes written into the EHR. Validated head to head against certified medical interpreters with physician-researchers from Johns Hopkins and the VA. (HealthTech, US)
Sending money home costs the most for those who can least afford it. SpotPay builds cross-border payments and savings for everyone. (Fintech)
Selling into healthcare is slow because the underlying data is broken. Paratus Health (YC W25) is building a go-to-market data platform for healthcare, from a team that previously scaled voice agents to 35+ outpatient facilities in under a year. (Healthcare IT, San Francisco)
Compute can’t get built because it can’t get power. Voxel Energy builds solar-plus-storage compute using repurposed EV battery packs, sidestepping the interconnection queue and giving second life to used packs. (Deep tech)
Investors with assets still borrow at over 18% because nobody offered them the alternative. Yenmo (YC W24) lets Indian investors pledge mutual funds and stocks digitally and draw a loan at a fixed 10.5% in under five minutes, instead of liquidating or taking a personal loan. (Fintech, Bengaluru)
Factories have decades of process knowledge and no way to use it. Allus AI puts an edge agent on the floor, so every install compounds the dataset. (Industrial AI)
Good financial advice is priced for people who already have money. Portfolio Pilot builds autonomous financial advice for individuals, in a regulated, high-error-cost domain. (Fintech)
Half of first attempts at a central line fail because ultrasound is still 2D. Lumius (YC P26) builds real-time 3D ultrasound that’s fast, portable and affordable. A Duke spinout. (MedTech, US)
What we’re looking for
If the last cycle rewarded generality, this one rewards teams who picked one workflow, wired it into the system of record, and can show the number moving. We’re actively meeting founders in five areas.
Robotics niches, not general-purpose humanoids. The general-purpose robot is a decade-long bet with a decade-long balance sheet behind it. The returns in this cycle sit in narrow, unglamorous jobs where the task is bounded, the environment is controlled and the payback is countable: sortation and tote handling, lab automation, agricultural picking, food assembly, inspection in places people shouldn’t be. Alongside that, the picks-and-shovels layer stays underpriced. Actuators, dexterous hands and end-effectors, fleet orchestration, safety certification, evaluation and simulation infrastructure. Everyone is buying the robot. Almost nobody is funding what the robot needs.
Egocentric data collection. This is the constraint we’re most interested in right now. Teleoperation is the expensive way to teach a robot: it needs hardware, expert operators and controlled setups, which is why real-robot data stays scarce. Humans, meanwhile, perform manipulation tasks all day for free. A head-mounted camera captures them from exactly the viewpoint the robot will occupy, including hand pose, object contact and gaze, none of which an external camera sees. The research has moved fast here. EgoVerse, released in April 2026 by researchers across Georgia Tech, Stanford, UC San Diego, ETH Zürich, MIT, Meta Reality Labs and Scale AI, is built as a living pipeline rather than a static dataset, and shows robot task performance improving logarithmically with hours of human demonstration. EgoDex reached 338,000 demonstrations and 829 hours of first-person video with full two-hand joint annotation. What’s still missing is the commercial layer: capture hardware people will actually wear, annotation pipelines that survive contact with a factory floor, and the rights and consent architecture to collect this data legally at scale. The embodiment gap between a human hand and a robot gripper remains an open research problem, which is precisely why this is an early-stage question rather than a late one.
And on the LP side
We’re finishing the build of our LP group, and the profile we keep coming back to isn’t defined by cheque size.
Operators from the sectors we invest in. Health system executives, plant and manufacturing leaders, energy and grid people, payments and lending operators. They price our diligence better than we do and open doors that cold outreach cannot.
Families and institutions with a genuine Transpacific footprint. Hong Kong, Singapore, Japan, Korea, Australia, and the US corridors we already work across. Our companies need distribution on both sides of the ocean, not just capital on one.
Founders and exited operators who want to stay close to the work. The people who answer a portfolio company’s message on a Sunday.
If that sounds like you, or someone you know, we’d welcome the conversation. And if you’re a founder building in any of the five areas above, reply to this email.
Sources: METR time-horizon research and Kwa et al.; Gartner and PwC agentic adoption forecasts; Salesforce customer-reported figures; Omdia and Counterpoint humanoid shipment estimates; NVIDIA, Physical Intelligence, Google DeepMind and Figure model disclosures; Agility Robotics and GXO deployment reporting; Lawrence Berkeley Lab, Dominion and FERC interconnection data; Drug Target Review and Bessemer Venture Partners. Third-party estimates vary by methodology and are presented as published evidence, not verified fact.
This newsletter is for information only. It is not an offer to sell securities, a solicitation of an offer to buy securities, or investment advice. Portfolio companies are named as illustrations of thesis, not as performance claims. Past performance is not indicative of future results.

