Scope at a glance
How to read this publication
- Domains
- **Domains** Retail and catalog operations, employment-related candidate workflows, and support operations.
- Economics
- **Economics** Cost per accepted output, quality-adjusted throughput at a fixed budget, oversight and rework time, escalation burden, and downstream outcome measures.
- Guardrails
- **Guardrails** Defect and constraint-violation rates, human escalation and override, fairness and impact monitoring where people are affected, data-protection posture, incident response, and rollback criteria.
Evidence taxonomy
Different claims carry different weight
External reported result
A company, vendor, or engineering team reported the result publicly; it has not necessarily been independently replicated.
Legal requirement
A requirement from a named law, rule, or official authority, scoped to the relevant jurisdiction and use.
Established method
A statistical or risk-management method from an authoritative technical source.
inAi proposed gate
A design, evaluation, or operating rule proposed by inAi for this framework.
Illustrative scenario
A synthetic example used to explain a calculation or monitoring pattern; not observed operational data.
inAi measured result
Reserved for documented inAi experiments, products, or partner deployments.
Operational research overview
Abstract and evidence boundary
AI is moving from isolated assistance into operational workflows, but the appropriate level of autonomy remains task-specific. This overview compares public case evidence across retail/catalog, employment-related candidate workflows, and support, then proposes acceptance gates, sampling, monitoring, rollback, and data-protection practices that organizations can adapt to their own risk tolerance and operating conditions. Reported company results, legal requirements, established methods, illustrative examples, and inAi recommendations are labelled separately.
Evidence and provenance legend
Evidence types are separated so that reported cases, legal requirements, established methods, illustrative examples, and inAi findings are not mistaken for one another.
Reversible configurations
Assist, supervise, and run as operating modes
Treat assist, supervise, and run as reversible operating configurations, not as a universal maturity ladder. A bounded task should move toward greater autonomy only when quality-adjusted value improves, severe-risk thresholds remain satisfied, human and exception burden remain acceptable, recovery is tested, and the legal and data posture fits the actual use. A task may remain supervised indefinitely, split into subtasks with different modes, or move back toward greater review when the model, data, environment, policy, or user impact changes.
Reversible operating configurations
Human primary
Sampled review + exceptions
Bounded intents + recovery
← keep · revise · split · revert →
Assist, supervise, and run are reversible operating configurations. The correct mode depends on the task, evidence, impact, and recovery conditions—not on a universal progression toward maximum autonomy.
Cases and boundaries
Task evidence by domain
| Task | Reported or proposed operating mode | Evidence / acceptance criteria | Reported operational result | Evidence status and date |
|---|---|---|---|---|
| **Retail — simple attribute extraction** | Automated production extraction with periodic sampling and low-confidence human review | Use attribute-specific human-labelled validation. Instacart reports **95% accuracy** for one simple “organic” attribute example; this is not a universal precision/recall threshold. LLM evaluation should be calibrated against human labels. | For simple attributes, Instacart reports similar quality from a less powerful model at **70% lower model cost**; for difficult attributes, the less powerful model produced a **60% accuracy drop**. | **External reported result** · Instacart engineering article · Aug 2025 |
| **Retail — complex or multimodal attribute extraction** | Automated extraction with periodic sampling and exception review | Validate exactness for numeric attributes, source consistency, missing evidence, and attribute-specific failure modes using a human-labelled sample. | For `sheet_count`, Instacart reports that multimodal LLMs increased recall by **10% over text-only models**. Preserve the source’s wording; do not silently change this to 10 percentage points. | **External reported result** · Instacart engineering article · Aug 2025 |
| **Retail — catalog ingestion, enrichment, and publication** | Automated infrastructure with validators, versioning, regression detection, and rollback | Define schema and business-rule validation, source traceability, catalog-level regression thresholds, and rollback tests. The Uber source does not report “zero high-severity violations” or WECO control rules. | Uber reports INCA processing **100,000 changes per second**, billions of changes daily, validation during publishing, regression detection, and near-instant rollback through catalog versioning. | **External reported production architecture** · Uber Engineering · Aug 2025 |
| **Employment-related candidate screening** | Human-accountable decision support; the appropriate mode depends on intended use, jurisdiction, and impact | Where NYC Local Law 144 applies: independent bias audit within one year before use, public summary, and required notice. Also evaluate job relevance, data quality, accessibility, human review, contestability, and other applicable anti-discrimination/privacy duties. | Do not claim that compliance “bounds legal risk.” Local Law 144 creates specific audit and notice duties; it does not eliminate broader legal or ethical risk. | **Legal requirement plus inAi proposed controls** · NYC DCWP materials · rechecked Jul 2026 |
| **Support — engineering ticket assignment** | Automated assignment with a locally defined exception or confidence path | Measure correct team assignment on a representative labelled sample, plus misrouting cost, exception rate, and time to correction. Confidence-gated escalation is an inAi recommendation unless separately documented by the source. | Google Cloud’s Gelato customer story reports accuracy increasing from **60% to 90%** and **120 hours of weekly labor saved** for engineering-support ticket assignment. | **Provider-hosted customer case study** · current source reviewed Jul 2026 |
| **Support — guarded knowledge-base Q&A** | Automated response with online guardrails, retry or escalation, offline quality monitoring, and human calibration | Use a shallow check first where appropriate, then an LLM evaluator for flagged cases; calibrate automated monitoring against random human-reviewed samples and track groundedness, accuracy, language, coherence, relevance, and severe-policy failures separately. | DoorDash reports a **90% reduction in overall hallucinations** and a **99% reduction in potentially severe compliance issues**. The source does not claim zero failures. | **External reported result** · DoorDash Engineering · Sep 2024 |
| **Support — end-to-end chat with human handoff** | Automated service for bounded intents with accessible human escalation | Track customer satisfaction, first-contact resolution, repeat inquiry rate, escalation, time to resolution, complaint rate, vulnerable-user outcomes, and service quality—not only containment. | Klarna reported in Feb 2024 that the assistant handled two-thirds of chats, reduced average resolution time from 11 minutes to under 2, and matched human-agent satisfaction. Later 2025 reporting described a strategic rebalancing toward human service and quality. | **Company launch report plus later independent update** · Feb 2024 / Sep 2025 |
Reported results, legal requirements, proposed gates, and evidence dates remain visibly separated.
Source-specific corrections
What each operational case supports—and where it stops
Instacart PARSE
- Attribute-specific thresholds on a human-labelled validation sample; Instacart reports 95% accuracy for one simple “organic” attribute example.
- LLM evaluation calibrated against human-labelled samples, with low-confidence and sampled outputs reviewed by humans.
- Instacart reports similar quality at 70% lower model cost for simple attributes; difficult attributes experienced a 60% accuracy drop with the less powerful model.
- Instacart reports a 10% recall increase over text-only models for the
sheet_countexperiment.
Uber INCA
Gate: schema and business-rule validation, source traceability, regression detection, versioned rollback, and incident-specific stop conditions.
Evidence status: externally reported production architecture; not an inAi autonomy measurement.
NYC Local Law 144
- These controls address specific Local Law 144 audit and notice requirements where the law applies; they do not eliminate wider discrimination, privacy, accessibility, employment, or data-protection obligations.
- Where Local Law 144 applies, preserve the independent bias-audit record, public summary, and required candidate or employee notices. Treat those requirements as a minimum, not a complete fairness or legal-risk framework. Ongoing monitoring should also examine job relevance, data quality, subgroup outcomes, human override, accessibility, contestability, and the consequences of false positives and false negatives.
Monitoring fields
- the monitored quantity;
- the denominator;
- the population and time window;
- minimum sample sizes;
- intersectional groups;
- what happens when small samples make the chart unstable;
- the difference between a legal audit, ongoing fairness monitoring, and a statistical process-control alert.
Gelato / Google Cloud
Google Cloud’s Gelato customer story reports that automated engineering-support ticket assignment increased correct routing from 60% to 90% and saved 120 hours of weekly labor. The source is a provider-hosted customer case study; local deployment should separately measure confidence, exception handling, misrouting cost, and time to correction.
DoorDash
- DoorDash reports a 99% reduction in potentially severe compliance issues; track the remaining severe-event rate, denominator, observation period, and escalation outcome.
- A shallow online check handles straightforward cases; flagged responses receive LLM-based evaluation, while broader offline quality monitoring is calibrated against random human-reviewed samples.
- LLM-evaluator requirement
Treat an LLM evaluator as a calibrated measurement instrument, not as ground truth. Validate it against a human-labelled set, report agreement by defect class and language, test order and prompt sensitivity, repeat a sample across runs, and route uncertain or high-impact cases to human review. Severe-policy gates should not depend on an uncalibrated reference-free judge alone.
Klarna
Klarna’s February 2024 launch metrics illustrate the scale and speed possible in bounded customer-service workflows. Later 2025 reporting indicates that the company rebalanced toward human availability, service quality, and growth. The combined case supports a reversible configuration model: automation levels should respond to customer outcomes and organizational strategy, not only containment and cost.
Outcome metrics
- repeat-contact rate;
- complaint rate;
- customer effort;
- successful human escalation;
- vulnerable-user outcome;
- abandonment;
- reversal or correction rate;
- long-tail intent coverage;
- employee knowledge and quality feedback;
- downstream revenue or retention where causally interpretable.
Illustrative scenario
Quality-adjusted unit economics
Cost per accepted output = (model + retrieval + guardrails + human review + rework + infrastructure + incident allowance) ÷ accepted outputs after the defined quality gate
Quality-adjusted throughput = accepted outputs without severe defect ÷ total human and machine operating time
Cost per accepted output
Quality gate · sample period · accepted count · severe-defect count
External case metrics provide context about scale and operating outcomes; they are not the source of the illustrative euro values.
Cost categories that must remain visible
- model inference;
- retrieval or data access;
- guardrails and evaluation;
- human review;
- rework;
- integration and orchestration;
- monitoring and observability;
- incident investigation and recovery;
- vendor and infrastructure overhead;
- legal/compliance operations where genuinely attributable;
- change management and worker training;
- downstream error cost.
Established methods
Sampling and evaluator calibration
Use classical attribute acceptance-sampling methods, such as those described in the NIST/SEMATECH Engineering Statistics Handbook, when the workflow genuinely forms inspectable lots and pass/fail defects can be defined. Select the sample size and acceptance number from explicit acceptable-quality, rejectable-quality, producer-risk, and consumer-risk assumptions.
Sampling assumptions and risk stratification
- a defined lot or time window;
- a stable defect taxonomy;
- random or representative sampling;
- clear handling of duplicate/correlated outputs;
- stratification where risk differs by language, category, retailer, candidate group, intent, or complexity;
- explicit producer and consumer risk;
- a procedure for rejected lots;
- a rule for increasing sampling after incidents or changes.
Sample proportionally to traffic for overall quality estimation, but oversample low-volume, high-impact, newly changed, multilingual, or historically weak segments. Report weighted aggregate results and segment-level results separately.
Illustrative monitoring
Drift, control limits, and recovery
Illustrative monitoring example — synthetic data
Illustrative monitoring example — synthetic data
This chart demonstrates how a quality signal, rollback, and increased sampling may be displayed. It is not a report of an inAi, PageMind, Emplo, customer, or partner incident.
A control chart detects changes in process behavior; it does not decide whether the baseline level is acceptable. Maintain separate quality, safety, legal, and policy thresholds.
Action-plan ownership
- who receives the alert;
- who can stop the workflow;
- who approves restoration;
- what evidence is needed;
- how affected users or partners are informed where appropriate.
Quality-adjusted operating loop
Economics · human impact · data and law · recovery
Operational AI is an iterative control problem. Deployment mode should change when evidence, risk, data, law, or downstream outcomes change.
Configuration-specific posture
Data protection and regulatory posture
Data protection and regulatory posture
Data location and retention must be verified for the exact provider, model, feature, account, region, and contract used by the workflow. For EU workloads, document where data is stored and processed, and check exceptions introduced by grounding, retrieval, agent features, session state, caches, request-response logging, abuse monitoring, and subprocessors. Do not claim zero data retention unless the complete configured data flow and current contractual terms support that claim.
Where the EU AI Act or another AI-specific regime may apply, classify the system by its intended purpose and distinguish the obligations of providers, deployers, employers, and other actors. Employment-related AI can fall into a high-risk category, but obligations depend on the use, role, timeline, and final applicable guidance. Keep this legal and regulatory note date-stamped and recheck it before publication or deployment.
Product workflows should define only the records they genuinely need: source and evidence links where relevant, action and approval logs, model and prompt versions, data-location settings, retention/deletion rules, incident records, and user or worker notices where required. PageMind, Emplo, and other products require separate product-specific assessments; this overview does not make a blanket compliance claim for them.
Task-specific controls
Domain playbooks
Retail and catalog workflows
- Use domain terminology controls and row-level source or evidence links where they improve reviewability. For EU workloads, verify the exact provider and feature configuration for data location, retention, logging, grounding, caching, and subprocessors; do not rely on an endpoint label alone.
- Measure downstream catalog consequences—rejection, correction, return, search failure, compliance escalation, and customer-impacting defect—not only extraction accuracy.
Employment-related candidate workflows
- Where an AEDT law or high-risk AI regime applies, preserve the required independent audit, public summary, notice, and role-specific records. Treat those controls as a minimum rather than a complete fairness or employment-law framework.
- Monitor selection or scoring outcomes, false-positive and false-negative consequences, job relevance, data quality, subgroup and intersectional effects, human override, accessibility, and contestability. Statistical stability does not by itself establish fairness.
- Preserve enough information to reconstruct the decision-support process: tool and version, intended use, inputs and features where lawful, output, human reviewer action, override or appeal, and the reason for the final decision. Avoid logging sensitive data merely because it could be useful later.
- Employment-related AI should support accountable human decision-making unless a lawful, validated, and appropriately governed use justifies another configuration. The page should not suggest that a bias dashboard alone makes automated shortlisting safe.
Support operations
- Use inexpensive deterministic or similarity checks for straightforward cases where they are valid, then apply an LLM evaluator, retry, or human escalation to flagged cases. Calibrate the evaluator against random human-reviewed samples and monitor severe defects separately.
- Track containment together with repeat contact, successful resolution, complaint rate, human-handoff success, customer effort, abandonment, and long-tail failure. A high automation rate is not sufficient evidence of a better service.
Engineering pattern
Migration checklist
- Define the task and operating unit. Specify the exact subtask, user, consequence, accepted-output definition, and conditions under which the system must abstain or escalate.
- Measure the baseline. Record current quality, throughput, human time, rework, escalation, downstream outcomes, and total operating cost before changing the mode.
- Choose task-specific acceptance criteria. Set quality, severe-defect, human-impact, legal, and recovery gates independently from the economic target.
- Validate on representative and risk-stratified data. Include difficult, multilingual, newly changed, low-volume, and high-impact segments.
- Calibrate automated evaluators. Compare LLM judges, confidence scores, and guardrails against human-labelled samples and define uncertainty or escalation behavior.
- Use sampling for the right purpose. Apply acceptance sampling for lot decisions where appropriate, and separate ongoing process monitoring for drift and diagnosis.
- Test failure and recovery. Verify rollback, quarantine, retry, human takeover, incident ownership, and restoration criteria before increasing autonomy.
- Verify data and legal posture. Check the actual provider, model, feature, region, logging, retention, contract, intended purpose, and organizational role.
- Increase autonomy gradually and reversibly. Start with a bounded scope or traffic slice, monitor outcomes, and be prepared to move back toward supervision.
- Review after material change. Re-evaluate the mode after changes to the model, prompt, tools, data, workflow, law, policy, or affected population.
This is longer than the current five points, but it is still compact and substantially more defensible.
Measurement architecture
Metric families and precise terminology
Metric families
15.1 Output quality
- precision, recall, accuracy, or exact match where appropriate;
- groundedness or source consistency;
- severe-defect rate;
- calibration;
- coverage and abstention;
- multilingual and segment-level performance;
- consistency across model or prompt versions.
15.2 Operational value
- accepted outputs per hour or fixed budget;
- cost per accepted output;
- review and rework time;
- escalation burden;
- latency to accepted result;
- incident and recovery cost;
- integration and maintenance burden.
15.3 Downstream outcomes
- catalog correction and rejection;
- customer return or complaint;
- repeat support contact;
- time to actual resolution;
- human-handoff success;
- candidate false-positive/false-negative consequences;
- employee or user outcome;
- revenue, retention, or risk only where causal interpretation is credible.
15.4 Human and organizational effects
- expertise retention and learning;
- reviewer fatigue;
- deskilling or overreliance;
- override and appeal behavior;
- distribution of gains across experienced and less-experienced workers;
- job redesign and role clarity;
- employee acceptance and work quality;
- ability to recover when the automated system is unavailable.
15.5 Governance and recovery
- audit-log completeness;
- incident detection time;
- time to containment;
- rollback success;
- restoration approval;
- affected population;
- notification and redress where relevant;
- provider/model/configuration inventory;
- date of last revalidation.
| Current term or pattern | Recommended replacement | Reason |
|---|---|---|
| candidate ops | employment-related candidate workflows | More precise and less dehumanizing |
| accepted item | accepted output | Works across all domains |
| autonomy migration | operating-mode selection or reconfiguration | Avoids implying inevitable one-way progress |
| move up the ladder | move toward greater autonomy / change mode | Reversible and task-specific |
| Auto | automated within defined scope | Avoids ambiguity |
| Production | externally reported production case / legal requirement / illustrative example | Reveals evidence type |
| LLM-judge pass | calibrated LLM evaluation against human labels | Judge is not ground truth |
| zero failures | observed severe-defect rate over N cases and period | Zero requires denominator and confidence interval |
| legal risk bounded | addresses named legal requirements; wider risk remains | Correct legal scope |
| zero retention | verified provider/model/feature-specific retention posture | Avoid blanket claim |
| compliance | named requirement or policy category | “Compliance” is too broad |
| constraint violation | defined defect or policy category | Requires a taxonomy |
| data protection | data location, retention, access, logging, deletion, and role-specific obligations | More operationally specific |
Glossary terms
- accepted output;
- operating mode;
- acceptance gate;
- severe defect;
- LLM evaluator;
- acceptance sampling;
- p-chart/u-chart;
- control limit;
- policy/specification limit;
- rollback;
- intended purpose;
- provider/deployer.
Non-claims
Limits of this overview
What current organizational research adds
Current field evidence does not support a single answer to whether AI should assist, supervise, or run work. Results vary by task, worker, expertise, workflow design, and how well the task fits the system’s capabilities. Some studies find substantial short-term productivity gains, especially for less-experienced workers or tasks inside the capability frontier. Other results show quality losses when users rely on AI outside that frontier. Research on teams also suggests that the design of human–AI collaboration and expertise integration matters, not only individual speed.
For business operations, the practical conclusion is to evaluate the complete work system: task decomposition, quality, human knowledge, escalation, incentives, downstream effects, and reversibility. Faster output is useful only when the organization can still detect error, preserve expertise, and achieve the intended outcome.
Evidence library
References by source type
Open reference library
Current page
Applied cases
- Instacart Engineering, “Scaling Catalog Attribute Extraction with Multi-modal LLMs”, 1 Aug 2025.
- Uber Engineering, “From Restaurants to Retail: Scaling Uber Eats for Everything”, 6 Aug 2025.
- DoorDash Engineering, “Path to high-quality LLM-based Dasher support automation”, 17 Sep 2024.
- Google Cloud, Gelato customer case study.
- Klarna, “AI assistant handles two-thirds of customer service chats in its first month”, 27 Feb 2024.
- OpenAI, Klarna case page.
- Reuters, “Europe’s AI poster child Klarna taps the brakes on chatbots”, 10 Sep 2025.
Employment and legal sources
- NYC Department of Consumer and Worker Protection, Automated Employment Decision Tools.
- NYC DCWP, AEDT Frequently Asked Questions.
- NYC Council, Local Law 144 legislative record.
- European Commission, Guidelines for providers and deployers of high-risk systems, current page reviewed Jul 2026.
- European Commission, Navigating the AI Act.
- European Commission, Draft high-risk classification guidelines, May 2026.
Data location and retention
- Google Cloud, Agent Platform and zero data retention.
- Google Cloud, Abuse monitoring.
- Google Cloud, Data residency.
- Google Cloud, Service Specific Terms.
- Google Cloud, Platform Terms of Service.
Statistical methods and AI risk management
- NIST/SEMATECH, e-Handbook of Statistical Methods.
- NIST/SEMATECH, What is Acceptance Sampling?.
- NIST/SEMATECH, Lot Acceptance Sampling Plans.
- NIST/SEMATECH, Choosing a Sampling Plan with a Given OC Curve.
- NIST/SEMATECH, Attribute Control Charts.
- NIST/SEMATECH, WECO Rules.
- NIST, AI Risk Management Framework.
- NIST, Generative AI Profile.
Organizational and productivity research
- Brynjolfsson, Li, and Raymond, “Generative AI at Work”, Quarterly Journal of Economics, 2025.
- Dell’Acqua et al., “Navigating the Jagged Technological Frontier”, Organization Science, 2026.
- Dell’Acqua et al., “The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork”, Organization Science, 2026.
- OECD, “The Effects of Generative AI on Productivity, Innovation and Entrepreneurship”, 2025.
- Raisch et al., “Roles of Artificial Intelligence in Collaboration with Humans”, Management Science, 2025.
LLM evaluation research
- Shi et al., “A Systematic Study of Position Bias in LLM-as-a-Judge”, 2025.
- Panickssery et al., “LLM Evaluators Recognize and Favor Their Own Generations”, NeurIPS 2024.
- Sheng et al., “Analyzing Uncertainty of LLM-as-a-Judge”, EMNLP 2025.
- Li et al., “Opportunities and Challenges of LLM-as-a-judge”, EMNLP 2025.
Source-card format and minimum set
SOURCE
Instacart Engineering — “Scaling Catalog Attribute Extraction with Multi-modal LLMs”
Published: 1 Aug 2025
Evidence type: company engineering report
Relevant scope: catalog attribute extraction
Reported result: 95% accuracy for one simple “organic” attribute; 10% recall increase for multimodal sheet_count; 70% lower model cost for simple attributes; 60% accuracy drop on difficult attributes with weaker model
Limits: not an independent cross-company benchmark; results are attribute- and configuration-specific
[Read source]
- Instacart PARSE.
- Uber INCA.
- DoorDash support automation.
- Gelato customer case.
- Klarna launch case plus later update.
- NYC Local Law 144 official materials.
- EU AI Act official materials.
- NIST acceptance sampling and control charts.
- NIST AI RMF.
This can be a compact references drawer rather than a large visible section.


