Radar Methodology
How we structure our market monitoring approach
Two main axes
Maturity (vertical axis)
This axis measures the depth and robustness of a solution today across three dimensions:
- Capabilities: Can the solution actually plan, reason, and adapt autonomously? Does it go beyond conversational interaction to execute actions in dynamic environments?
- Reliability: Is the solution stable in production? Can it scale from 10 to 10,000 users and remain resilient when an external API goes down? Are errors handled gracefully rather than triggering cascading failures?
- Trust: Can the solution be deployed in a regulated environment? Are the agent’s decisions explainable? Is data secure? Are guardrails in place to prevent undesirable behavior?
Momentum (horizontal axis)
This axis measures a solution’s velocity, its ability to engage a user community, gain enterprise adoption, and earn market recognition.
- Velocity: What is the release cadence? How strong are technical support and documentation? How frequent and substantial are contributions to public repositories?
- Engagement: Is there an active community of contributing developers? How many third-party integrations are available? How strong are the plugin and tooling ecosystems?
- Traction: How many customers are using the solution in real production environments, not just POCs? What strategic partnerships has it established? How is it recognized by independent analysts?
Four main zones
Bottom left, the Innovators: these solutions sit at the boundary between academic research and early experimental implementations. They show a high pace of publishing and experimentation, including frequent commits, arXiv papers, and public proofs of concept, but their production maturity remains low: unstable APIs, incomplete documentation, and error handling that is not yet sufficient for real-world deployment. These are the solutions that will shape tomorrow’s Builders and Drivers.
Lower middle zone, the Builders: solutions in the Builder zone have crossed the threshold of technical viability. But not all solutions under development follow the same path. The Radar distinguishes between two subprofiles with very different implications.
- Maturity Builders have momentum but invest primarily in depth. They strengthen their internal architecture, improve error handling, expand technical documentation, and build the foundations required for production readiness. Their communities are smaller but more specialized, with core contributors and enterprise architects.
- Momentum Builders, by contrast, capitalize on rapid adoption: multiple integrations, a growing plugin ecosystem, rapidly increasing star counts, and frequent releases. However, each release may introduce breaking changes, while managing technical debt takes a back seat to user acquisition.
Upper middle zone, the Drivers: these solutions have reached the critical threshold where maturity and momentum reinforce one another. Here again, the Radar distinguishes between two strategically different subprofiles.
- Maturity Drivers have built their position through maturity. They are historically robust, battle-tested in critical production environments, and supported by conservative release processes that preserve stability but slow the adoption of new capabilities. Their momentum is real but measured.
- Momentum Drivers have built their position through adoption and are consolidating it through maturity. They have broad communities and rich ecosystems, while successfully crossing the production-readiness threshold without losing velocity.
Top right, the Leaders: these solutions combine proven production maturity at scale with enough momentum to maintain their technological lead. They can meet the demands of critical systems, including scalability, resilience, auditability, and regulatory compliance, while continuing to integrate new capabilities at the pace of the market. They are established choices for enterprise technology portfolios.
The evaluation process
Four pillars
This Radar is based on in-depth AI-assisted research backed by human expertise: technical documentation, Discord and GitHub discussions, scientific publications, and commit tracking. AI accelerates market monitoring by parsing thousands of forum messages, extracting trends, and identifying weak signals. Academic publications, particularly on arXiv, are a valuable source of early insight, revealing the architectures that may shape the frameworks of the future.
We complement this research with field feedback from our experts. When frameworks such as LangGraph or Strands Agents have been implemented for clients, architects and developers bring irreplaceable hands-on experience: what actually works, what is difficult to maintain, and where the pitfalls are. This field feedback helps rebalance assessments that might otherwise be influenced by vendor marketing.
Finally, we run benchmarks using standardized scenarios: the same task, the same constraints, and the same allocated resources. This comparative rigor is particularly important because some models have demonstrated an ability to recognize when they are being evaluated. One documented case involved a model explicitly identifying that it was undergoing benchmark testing after finding the test protocol online.
Integrity principles
Two principles govern the process. First, the exact scoring criteria are not published. In a world where agentic solutions can themselves access online benchmarks, publishing the exact criteria would introduce bias by encouraging solutions to optimize for the evaluation rather than demonstrate genuine maturity.
Second, the Radar is developed, updated, and powered by a fully internal tool called Elio, an agent platform built for our teams. This removes preference bias and conflicts of interest.
These principles draw directly from established practices in independent evaluation used in other industries, such as financial auditing and pharmaceutical testing. A framework evaluating other frameworks while using one of those frameworks as the evaluation tool would introduce a structural bias that would be difficult to quantify and impossible to correct after the fact.
Solution typology
To distinguish the technical nuances among the different agentic solutions available, we use these categories:
-
Dev / Test / Tech: Solutions that help technical teams design, develop, test, debug, evaluate, and deploy AI applications or agents.
-
Orchestration & Automation: Solutions that coordinate tasks and agents, apply business rules, and automate complex workflows across multiple systems.
-
Infra, Runtime & Ops: Solutions that provide the infrastructure needed to deploy, run, monitor, and optimize agents in production.
-
Business Applications & Suites: Software suites that embed agents directly into existing business applications and processes, including CRM, ERP, productivity, and HR platforms.
-
Conversational & Experience Platforms: Platforms that enable organizations to build and deploy user-facing agents across channels such as chat, web, and voice, including for support and workplace interfaces.
-
Analytics, Decision & Knowledge: Solutions that connect agents to enterprise data, documents, knowledge, and analytics tools to support or automate decision-making.
-
Governance, Risk & Compliance (GRC): Solutions that govern the use of agents through security controls, access management, auditing, compliance, risk management, and oversight.
-
Industry-Specific Solutions: Solutions designed to address the specific needs and constraints of a given industry or specialized function, including models, data, workflows, and controls.
-
Cross-Domain Agent Platforms: General-purpose platforms that enable organizations to build, connect, deploy, and manage agents across multiple business functions.