← Tim Vasil

How AI-native is your technical organization?

By Tim Vasil •

Designed to replace hype and FOMO with empirical rigor, this calculator provides technical leaders with a holistic framework for estimating GenAI-enabled productivity gains across the PDLC. What this calculator doesEstimates how much time AI could save, including work that still needs human attention.Includes time spent reviewing AI output, fixing mistakes, and maintaining the tools.Adjusts for your team’s size, experience, work mix, platform maturity, and other responsibilities.Compares your current setup with your target, including the initial slowdown while the team adapts.Shows suggested autonomy ranges and flags choices that may need closer oversight.Links to supporting research and explains the assumptions under “How this is estimated.”Lets you save and share your settings with a link.

Sources

The strongest theme across firsthand accounts from SaaS operators and AI engineering teams: automate execution aggressively, while humans own intent, boundaries, and evidence of success. The disagreement is over how much implementation and review can already be delegated. Highlights mark each source’s key take-away, and the last column records what it actually measured, if anything.

SourceWhat it saysPractical implicationReported impact
Tobi Lütke & Farhan ThawarShopifyInterview, 2025 AI adoption needs organizational support: broad access, shared integrations, reusable workflows, and leadership participation. Provide common infrastructure while enabling teams to develop their own applications.
  • No measured organization-wide gain in this account.
Kief MorrisThoughtworksHumans and agents, 2026 Humans build and manage the delivery loop, defining outcomes and improving how agents produce software. Move engineering effort toward architecture, verification, and intervention where judgment matters.
  • Operating framework, not a productivity study.
OpenAI engineeringHarness engineeringAccount, Feb 2026 One team built a new product with agent-written code and predominantly agent-based review. Humans supplied intent and the execution environment. Invest in executable constraints, accessible context, isolated environments, and agent-visible UI and telemetry.
  • ≈10× development speed; ≈90% less time. The team’s estimate against writing the code manually; one greenfield product, not a controlled comparison.
Justin McCarthyStrongDMSoftware factory, Feb 2026 Specifications and independently held scenarios drive implementation without human code review. Simulated services enable extensive validation. Explore autonomous delivery where independent validation is strong enough to support it.
  • No comparable before-and-after speedup.
  • Token-spend targets are not productivity evidence.
Anthropic & OpenAIOversight and permissionsAutonomy research, Feb 2026 Agent guide, 2025 Effective oversight combines visibility, intervention, bounded permissions, and escalation for failures or consequential actions. Grant authority per workflow and risk level; make agents inspectable and interruptible.
  • Autonomy and oversight findings, not delivery-speed estimates.
AnthropicAgent evaluationsDemystifying evals, Jan 2026 Evaluate actual outcomes and state changes. Automated checks, production monitoring, and human assessment catch different failures. Give teams ownership of representative evals and continuously improve them from failures.
  • Reliability guidance; no general speedup estimate.
AnthropicOrchestrationEffective agents, Dec 2024 Multi-agent research, Jun 2025 Use simple workflows where possible; parallel agents help when work genuinely separates, with additional cost and coordination. Use multiple agents selectively and measure their incremental benefit.
  • Up to 90% less research time, from parallelizing previously sequential agent execution on complex queries. Not versus human-only work.
AnthropicInternal engineering studyHow AI is transforming work, 2025 Engineers report broader output and capability, while most still fully delegate only a minority of their work. Assess effective collaboration and verification, not just maximum autonomy.
  • +50% self-reported productivity, ≈1.5×, averaged across surveyed engineers and researchers. Internal survey; not experimentally established.
Cui et al.Three developer field experimentsResearch paper, 2025 Randomized access to a coding assistant across 4,867 developers increased completed tasks; gains varied across developers. Experimental evidence for AI-assisted development, without assuming wholesale workflow redesign.
  • +26.08% completed tasks, ≈1.26× throughput; standard error 10.3 percentage points. Coding assistance, not an AI-native organization.
Demirer, Musolff & YangWriting versus shippingWorking paper, 2026 Across successive tool generations, coding gains attenuate substantially before release. The study uses GitHub activity and AI telemetry. Model downstream bottlenecks; distinguish generating changes from shipping useful software.
  • +240% commits, ≈3.4×, cumulative adoption through autonomous agents.
  • +30% releases, ≈1.3×: the same adoption, measured at what ships. Observational estimates, not randomized results or team-wide speed measurements.
METRExperimental productivity evidenceStudy, Jul 2025 2026 update Early-2025 tools slowed experienced developers on familiar repositories. Later measurement became compromised by selection effects and concurrent agent work. Keep the negative result as historical evidence; measure current tools in your own environment.
  • 19% longer task completion in early 2025, ≈0.84× speed.
  • The 2026 experiment does not establish a reliable replacement estimate.
METRTechnical-worker surveySurvey, May 2026 Technical workers report substantial gains, but perceived speed and value differ, and estimates may be overstated. Ask separately about time saved, valuable output, and confidence in the estimate.
  • Median 3× self-reported speed.
  • 1.4–2× self-reported value, depending on the question. Convenience sample of 349 technical workers; not causal evidence.
DORAOrganizational conditions2025 report AI amplifies existing organizational strengths and weaknesses. Platforms, workflows, and team alignment influence outcomes. Use readiness and delivery constraints to qualify the calculator’s estimates.
  • No universal AI-native multiplier; reported relationships are context-dependent.
DORAROI of AI-assisted software developmentFramework and calculator, 2026 A framework for evaluating AI investment, including the initial productivity dip and the economics of adoption. Use its framework and companion calculator to structure costs, benefits, and assumptions.
  • No universal AI-native multiplier. Models potential ROI for the organization.
GitClearCode quality and productivity researchCode-quality report, 2025 Productivity research, 2026 Churn, duplication, and moved-code trends are maintainability signals. They do not establish that AI caused deterioration or directly measure software shipped. The 2026 study separately examines AI-use cohorts. Track durable changes, rework, duplication, and review burden alongside output. “Moved code” is a proxy for refactoring, not all refactoring.
  • No causal speedup estimate.
  • 4–10× output differences between usage cohorts in the 2026 study, which explicitly questions how much reflects who adopts AI.
JellyfishState of engineering managementSurvey of 636 leaders and practitioners, 2026 Only 10% reported strong enablement and high adoption. AI use varied substantially by activity: writing code 53%, review 49%, specifications 24%. Broad, institutionalized adoption remains less common than individual use. The survey’s categories do not map directly onto the profiles here.
  • Adoption rates, not a speedup: 10% with strong enablement and high adoption.