How AI-native is your technical organization?
By Tim Vasil •
Designed to replace hype and FOMO with empirical rigor, this calculator provides technical leaders with a holistic framework for estimating GenAI-enabled productivity gains across the PDLC. What this calculator doesEstimates how much time AI could save, including work that still needs human attention.Includes time spent reviewing AI output, fixing mistakes, and maintaining the tools.Adjusts for your team’s size, experience, work mix, platform maturity, and other responsibilities.Compares your current setup with your target, including the initial slowdown while the team adapts.Shows suggested autonomy ranges and flags choices that may need closer oversight.Links to supporting research and explains the assumptions under “How this is estimated.”Lets you save and share your settings with a link.
Sources
The strongest theme across firsthand accounts from SaaS operators and AI engineering teams: automate execution aggressively, while humans own intent, boundaries, and evidence of success. The disagreement is over how much implementation and review can already be delegated. Highlights mark each source’s key take-away, and the last column records what it actually measured, if anything.
| Source | What it says | Practical implication | Reported impact |
|---|---|---|---|
| Tobi Lütke & Farhan ThawarShopifyInterview, 2025 | AI adoption needs organizational support: broad access, shared integrations, reusable workflows, and leadership participation. | Provide common infrastructure while enabling teams to develop their own applications. |
|
| Kief MorrisThoughtworksHumans and agents, 2026 | Humans build and manage the delivery loop, defining outcomes and improving how agents produce software. | Move engineering effort toward architecture, verification, and intervention where judgment matters. |
|
| OpenAI engineeringHarness engineeringAccount, Feb 2026 | One team built a new product with agent-written code and predominantly agent-based review. Humans supplied intent and the execution environment. | Invest in executable constraints, accessible context, isolated environments, and agent-visible UI and telemetry. |
|
| Justin McCarthyStrongDMSoftware factory, Feb 2026 | Specifications and independently held scenarios drive implementation without human code review. Simulated services enable extensive validation. | Explore autonomous delivery where independent validation is strong enough to support it. |
|
| Anthropic & OpenAIOversight and permissionsAutonomy research, Feb 2026 Agent guide, 2025 | Effective oversight combines visibility, intervention, bounded permissions, and escalation for failures or consequential actions. | Grant authority per workflow and risk level; make agents inspectable and interruptible. |
|
| AnthropicAgent evaluationsDemystifying evals, Jan 2026 | Evaluate actual outcomes and state changes. Automated checks, production monitoring, and human assessment catch different failures. | Give teams ownership of representative evals and continuously improve them from failures. |
|
| AnthropicOrchestrationEffective agents, Dec 2024 Multi-agent research, Jun 2025 | Use simple workflows where possible; parallel agents help when work genuinely separates, with additional cost and coordination. | Use multiple agents selectively and measure their incremental benefit. |
|
| AnthropicInternal engineering studyHow AI is transforming work, 2025 | Engineers report broader output and capability, while most still fully delegate only a minority of their work. | Assess effective collaboration and verification, not just maximum autonomy. |
|
| Cui et al.Three developer field experimentsResearch paper, 2025 | Randomized access to a coding assistant across 4,867 developers increased completed tasks; gains varied across developers. | Experimental evidence for AI-assisted development, without assuming wholesale workflow redesign. |
|
| Demirer, Musolff & YangWriting versus shippingWorking paper, 2026 | Across successive tool generations, coding gains attenuate substantially before release. The study uses GitHub activity and AI telemetry. | Model downstream bottlenecks; distinguish generating changes from shipping useful software. |
|
| METRExperimental productivity evidenceStudy, Jul 2025 2026 update | Early-2025 tools slowed experienced developers on familiar repositories. Later measurement became compromised by selection effects and concurrent agent work. | Keep the negative result as historical evidence; measure current tools in your own environment. |
|
| METRTechnical-worker surveySurvey, May 2026 | Technical workers report substantial gains, but perceived speed and value differ, and estimates may be overstated. | Ask separately about time saved, valuable output, and confidence in the estimate. |
|
| DORAOrganizational conditions2025 report | AI amplifies existing organizational strengths and weaknesses. Platforms, workflows, and team alignment influence outcomes. | Use readiness and delivery constraints to qualify the calculator’s estimates. |
|
| DORAROI of AI-assisted software developmentFramework and calculator, 2026 | A framework for evaluating AI investment, including the initial productivity dip and the economics of adoption. | Use its framework and companion calculator to structure costs, benefits, and assumptions. |
|
| GitClearCode quality and productivity researchCode-quality report, 2025 Productivity research, 2026 | Churn, duplication, and moved-code trends are maintainability signals. They do not establish that AI caused deterioration or directly measure software shipped. The 2026 study separately examines AI-use cohorts. | Track durable changes, rework, duplication, and review burden alongside output. “Moved code” is a proxy for refactoring, not all refactoring. |
|
| JellyfishState of engineering managementSurvey of 636 leaders and practitioners, 2026 | Only 10% reported strong enablement and high adoption. AI use varied substantially by activity: writing code 53%, review 49%, specifications 24%. | Broad, institutionalized adoption remains less common than individual use. The survey’s categories do not map directly onto the profiles here. |
|