Know whether one workflow will pay before you fund the pilot.
A fast demo can still become an expensive failure. Count integration, review, exceptions, and adoption before approving the pilot.
Lower model prices leave workflow fit unresolved.
Model cost has moved. Measured returns remain uneven. Task fit still decides whether AI helps or hurts.
-
01
Epoch AI & Stanford HAI
Model use can be dramatically cheaper at the same measured capability.
Epoch AI estimates blended prices at GPT-4-level GPQA performance fell by about 40x per year. The trend normalizes capability and mixes input and output costs, which differs from a raw token-price comparison. Stanford separately reports a 280x drop at GPT-3.5-level MMLU performance from November 2022 to October 2024.Decision: re-price the workflow, including integration, human review, and change effort.
Data through 2025 -
02
S&P Global, May 30, 2025
AI initiatives often stop before production or broad adoption.
In S&P Global's survey of 1,006 midlevel and senior IT and line-of-business professionals in North America and Europe, 42% of companies abandoned the majority of AI initiatives before production, up from 17% a year earlier. On average, 46% of projects were scrapped between proof of concept and broad adoption. Organizations with lower failure rates were more likely to prioritize compliance, risk, and data availability.Decision: screen compliance, risk, and data availability before approving the pilot.
S&P Global research, May 30, 2025 -
03
Harvard Business School & BCG
Task fit can create gains, then reverse outside the boundary.
In a preregistered field experiment with 758 BCG consultants, 385 joined the inside-the-frontier task arm and 373 the outside arm. Inside the frontier, GPT-4 users completed 12.2% more tasks, worked 25.1% faster, and produced responses rated 32% higher on average. In the separate outside-the-frontier task, assisted participants were about 19 percentage points less likely to reach the correct solution.Decision: screen the task before the tool and define the pilot's permitted boundary.
Experiment run 2023 · published 2026
Six inputs expose the real business case.
Use a current process and actual operating data where available. The result should show what must be measured and remain neutral about the answer.
-
01
Volume: how many cases repeat?
Count cases by week or month, include seasonal peaks, and separate repeatable work from exceptions.
-
02
Manual hours: where does time actually go?
Measure handling, rework, coordination, waiting, and manager review instead of using one broad estimate.
-
03
Error or delay cost: what gets expensive?
Track correction effort, missed deadlines, queue delay, lost capacity, or delayed revenue using documented values.
-
04
Integration: what must connect?
List data sources, permissions, identity, the system of record, handover needs, and ongoing maintenance.
-
05
Human review: which decisions stay named?
Define review time, escalation rules, approval authority, and the cases that must stop for a person.
-
06
Adoption: who must change behavior?
Name the users, training, fallback path, usage measure, and owner responsible for sustained use.
Market evidence informs the test. Your baseline decides it.
The three facts above drive the test. These sources add global and Mongolian context.
These sources add global context rather than a Mongolia-specific return forecast. Mongolia's cabinet source records an announced plan, not a procurement mandate.