Limits of Intelligence
What can intelligent systems do reliably, where do they fail, and which limits can be detected, reduced, or repaired?
Limits of Intelligence is inAi’s research direction for the boundaries of current AI models and the systems built around them. We study reasoning and generalization, context and memory, uncertainty and hallucination, evaluation, and the ways tools, retrieval, verification, and orchestration can reduce some failures while creating others.
The title describes a long-term research horizon. Our current public work is narrower and more practical: it examines model-and-agent systems operating with incomplete information, finite time and compute, imperfect tools, and real consequences when they are wrong.
Intelligence is not only what a system can do at its best
A system may solve difficult problems and still fail on a small change in wording. It may accept a long context without using the relevant part, continue after its evidence runs out, or produce a plausible result without recognizing that it should stop. In a longer workflow, one weak decision can propagate through memory, tool calls, verification, and execution.
For this direction, capability is not only the best result a system can produce. It also includes the conditions under which that result remains reliable, how failure becomes visible, whether the system can recover, and what happens when the environment becomes more novel, complex, long-running, or uncertain.
This is why inAi studies useful intelligence under real constraints. Cost and latency are not intelligence by themselves, but they matter when a system must choose how much reasoning, retrieval, verification, or review a task can justify. A method that improves an answer while making the workflow too slow, expensive, opaque, or unstable may move the boundary rather than remove it.
The anatomy of a limit
A limit is not only the point at which a system fails. It also includes whether the failure can be detected, contained, repaired, and learned from.
The broad question and the current scope
Limits of Intelligence is intentionally broader than a single benchmark or technique. Over time, the direction can include theoretical and practical questions about what intelligent systems can learn, generalize, remember, decide, and accomplish.
The current program begins with the operational limits of modern model-and-agent systems. It concentrates on five connected areas.
Reasoning and generalization
How systems handle novel problems, interacting rules, ambiguity, and changes in task structure — and where fluent reasoning becomes brittle or dependent on familiar patterns.
Context, memory, and state
How systems find and use relevant context, preserve information across time, distinguish useful memory from noise, and recover when stored state is incomplete or wrong.
Uncertainty and hallucination
How systems behave when evidence is missing, conflicting, or insufficient; whether confidence tracks correctness; and when the right action is to answer, retrieve, ask, abstain, or stop.
Tools, retrieval, and orchestration
When decomposition, routing, retrieval, memory, verification, and tool use improve performance — and when they add coordination overhead, latency, cost, or new failure modes.
Evaluation, repair, and operational reliability
How to measure more than a single final answer: consistency across attempts, long-horizon completion, error propagation, repairability, traceability, robustness, latency, and cost.
These five areas do not claim to exhaust the limits of intelligence. They define the current starting point for work that inAi can describe and publish responsibly.
Questions guiding the current agenda
A research direction becomes meaningful when it is organized around questions that can be examined, challenged, and revised. The current agenda is:
Which failures belong to the base model, and which are created or amplified by the surrounding system?
When do decomposition, routing, retrieval, memory, verification, and additional test-time computation improve the result — and when do they merely move, hide, or compound the failure?
Can a system recognize that it lacks evidence and choose reliably among answering, retrieving, using a tool, asking for review, abstaining, or stopping?
How does reliability change as tasks become longer, more novel, more interactive, or more dependent on imperfect tools and stored state?
How should intelligent systems be evaluated when real usefulness depends on consistency, repair, traceability, latency, and cost as well as final-answer quality?
Which limits can system design reduce, and which remain unresolved even when more compute, tools, memory, or coordination are added?
These questions may lead to review papers, research notes, evaluation protocols, failure taxonomies, comparative studies, or experiments. An item will appear as a public output only when it exists and can be labelled accurately.
AGENDA REVIEWED: JULY 2026
Current public work
The collection currently contains one substantive public output. It addresses one important part of the direction, but it should not be read as a complete theory of intelligence or as a claim that orchestration always outperforms larger models.
Limits of Intelligence — orchestration vs raw capacity
This overview examines when step decomposition, uncertainty-aware routing, retrieval, memory, verifier layers, and plan–act–verify loops can improve cost, latency, stability, and auditability compared with relying on raw model scale alone. It also considers anti-patterns and open problems, including imperfect verifiers, repair budgets, multilingual calibration, and selective retrieval.
- Institutional author
- inAi
- Published / last updated
- 1 October 2025
- Evidence window
- 2024–October 2025
- Publication status
- Published research overview
This is an inAi research overview. It is not presented as an academically peer-reviewed paper.
Read the research overviewNo other outputs are listed at present. The collection should grow through real publications, not placeholders.
How work in this direction is prepared
Work in this direction may combine literature review, benchmark analysis, conceptual modelling, evaluation design, failure taxonomy, and applied observation. The method should follow the question, and each output should state its evidence window, sources, limitations, authorship, publication date, and review status.
We distinguish external findings from inAi interpretation and treat negative or contradictory evidence as part of the research. Product and Open Source work may expose useful questions and failure modes, but a product result is not automatically a general research finding. The AGI-as-a-system thesis is a frame to test and refine, not a conclusion the research is required to protect.
A review paper, research note, essay, experiment, protocol, or technical analysis should be labelled as what it actually is. Company review should not be described as academic peer review.
Selected foundations and further reading
These sources help define the wider field around this direction. They are not a substitute for the source list behind each inAi publication, and inclusion does not mean that inAi adopts every conclusion.
On the Measure of Intelligence — François Chollet, 2019
A framework for thinking about intelligence through generalization and skill-acquisition efficiency rather than task skill alone.
ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems — ARC Prize, 2025
A benchmark designed to stress abstract reasoning and generalization on novel tasks that remain accessible to human solvers.
Lost in the Middle: How Language Models Use Long Contexts — Liu et al., TACL, 2024
Evidence that accepting a long context does not guarantee reliable use of the information within it.
Why Language Models Hallucinate — Kalai, Nachum, Vempala, and Zhang, 2025
An analysis of why training and evaluation can reward guessing instead of acknowledging uncertainty.
Measuring AI Ability to Complete Long Tasks — METR, 2025
A method for evaluating AI capability through reliable completion of tasks of increasing duration.
Test-time compute and verifier limits — Snell et al., 2024; Yu, Li, and Wang, 2025
Complementary work on when additional inference and verifier-guided search improve performance, and where imperfect verification creates scaling limits.
Why this matters to AGI as a System
The broader AGI as a System stance is the frame around these relationships. Tools, memory, feedback, and coordination may reduce some model limits, but every additional layer creates dependencies and failure modes of its own.
Limits of Intelligence studies that trade-off. The purpose is not to prove that systems always outperform individual models. It is to understand when system design extends capability, when it only shifts the cost or failure elsewhere, and which boundaries remain.
Where the research matters in practice
For Products for Agents, this work informs questions about state, recovery, permissions, verification, and safe stopping. For Trust, it supports clearer distinctions between capability, uncertainty, maturity, and evidence. For AI for Everybody, it provides a technical foundation for explaining why fluent output is not the same as reliable intelligence.
Research collaboration
inAi is open to serious collaboration around evaluation protocols, replication, long-context and memory limits, uncertainty and calibration, verifier reliability, agent failure and recovery, and the operational measurement of AI systems.
Relevant work may include a jointly defined review, external critique of an existing output, a comparative evaluation, a failure taxonomy, a protocol, a replication study, or a scoped experiment. A useful proposal should identify the question, method, expected output, responsibilities, data or confidentiality needs, and whether the result is intended for public release.
The aim is not to create research for appearance. It is to produce work that makes the boundaries of intelligent systems clearer, more testable, and more useful.
