AI CodingIndustry & Competition

Agents Mentioned PayPal 139 Times and Picked It 0 Times: Software Buyers Are No Longer Human, and Distribution Rules Have Changed

In the software procurement chain, a new buyer has recently emerged: the coding agent. It directly writes code, invokes tools, and modifies repositories; which developer tools actually make it into a project depends entirely on what it picks.

In September 2026, startup Armature released a large batch of empirical test data that mapped out this new buyer’s behavior: running a total of 16.9K sandboxed sessions where three mainstream agents—Claude Code, Codex, and Cursor—actually integrated developer tools in synthetic repositories, rather than asking them for recommendations in a chat box.

The test results broke conventional wisdom. The highly recognizable PayPal was mentioned 139 times in the initial release batch, but was actually picked and written into the codebase 0 times. The very same tool that is universally recognized as the top choice in a repository of one tech stack is knocked out immediately when moved to another.

Whether you build developer tools yourself or use agents to write code every day, this data forces you to rethink the logic of software distribution: what kinds of tools actually get picked by agents, and how much utility remains in the traditional playbook built on brand and marketing. In that same month, Y Combinator announced a hackathon themed Make Something Agents Want, pointing the entire ecosystem’s attention in the exact same direction as this data.

As buyers shift from humans to agents, competition moves from the product form layer to the selection layer

What the Empirical Data Proves, and Where Its Extrapolations End

What this data can answer is how agents actually pick tools, with three core conclusions: being mentioned does not equal being adopted; tool selection follows the language environment; and large markets are highly concentrated. Specific figures for each conclusion will be provided below; before that, consider two things: how the data was generated, and whether it can be trusted. They directly determine how far these numbers can prove.

First, consider how the data was generated. Armature prepared dozens of synthetic repositories with tech stack distributions aligned with public GitHub repositories, containing fictitious enterprise backgrounds and real dependency lockfiles. The tasks required agents to complete authentic tool integrations: integrating payment functionality, or configuring transactional email delivery. Across the entire process, only two things were recorded: which tools agents mentioned, and which ones were actually written into the code. Each task was paired with four developer personas: the delivery-speed-focused vibe coder, junior, senior, and enterprise engineers.

The integrations were run by three mainstream agents: Claude Code, Codex, and Cursor, whose underlying models were Claude Opus 5, GPT-5.6 Sol, and Grok 4.6, respectively. Each session featured a simulated user (played by a Gemini instance, barred by rule from proactively naming specific products) interacting with the agent across multiple turns on behalf of a real developer; throughout the process, 811 plan clarifications and 734 rejections for modification were recorded in the interaction logs, and at session wrap-up, the simulated user approved 100% of the patches. Finally, another independent Gemini instance served as judge, reading through the interaction logs and final code patches to tally which tools were mentioned and which vendor ultimately won.

The test scope warrants a discount first. Out of the full 16.9K sessions, the initial release includes only 5.3K records up to September 2, 2026; the remaining ~69% were excluded, primarily for three reasons: agent version updates causing entire batches of legacy data to be quarantined, unvetted evaluation sets being rerun in full, and fixed truncation of the first N runs in a preset order during each round of testing to prevent cherry-picking. For specific processing rules, see Armature’s measurement methodology documentation.

Credibility also calls for caution. Armature’s primary business is providing technical growth services starting at $5,000 per month to help developer tools win adoption by coding agents, and the company disclosed this conflict of interest at the beginning of the report: “Disclaimer: Armature sells growth services to dev tools. This study is part of our broader work on how to influence coding agents choices and get products picked.” Given these commercial interests, that is precisely why we verified it: we downloaded Armature’s publicly released dataset agent-leaderboards.csv under the CC BY 4.0 license along with full API data, filtered it by the September 2 cutoff, and reproduced the statistical methodology across all 5.3K sessions bit by bit; the core figures in the report match the underlying raw records.

First, being mentioned does not equal being adopted. In payments, a total of 395 integration tests were conducted; the widely known PayPal received 139 organic mentions, but the number of times it actually entered the codebase was 0; in those 139 runs, Stripe won 124 times. The tools an agent mentions in conversation and the tools it writes into code are completely different things.

Second, tool selection follows the language environment. Faced with the exact same transactional email delivery requirement, repositories in four different programming languages selected four distinct winners: the TypeScript repository chose Resend (55/89), the Python repository chose SendGrid (22/24), the Go repository chose Postmark (20/24), and the Java repository chose Azure ACS (22/23). Four languages corresponded to four winners, each achieving a win rate of over 60% in its respective language. While the sample size of this test group is limited, the direction is clear: the language environment dominates tool selection.

Third, large markets are highly concentrated. In the sample, Stripe achieved a win rate as high as 88% (349/395) in payments; Neon was similarly dominant in databases, winning 236/356 for a 66.3% win rate, ranking first in adoption rate across independent statistics for all three tested agents. Looking across developer personas: under the junior engineer persona, Neon was selected 104/106 times, and under the delivery-speed-focused vibe coder persona, it swept 60/60 times.

Conversely, there are also two things this data cannot prove. First, win rate does not equal market share: the simulated user in the tests stripped away existing human brand preferences and legacy procurement contracts, and the test environments were brand-new greenfield projects; Armature itself acknowledged in the report that in the real world, as long as humans participate in decision-making, incumbent brands will still capture a share. Second, agent-initiated deployments do not equal autonomous procurement: according to data disclosed by Vercel, although agent-initiated deployments on its platform exceed 50%, that is merely an action at the execution layer, not a selection decision—the final authority over budget approvals and account linkage remains firmly in the hands of engineers.

Who ranks first in the sample is therefore merely a byproduct of this data. Its real value lies in offering the industry a reproducible observation window for the first time, allowing us to see clearly how agents actually make tool selection decisions. The rest of this article uses this window to answer three core questions: where the competition in software distribution is taking place today; which analogies built on traditional search still hold and which have completely broken down; and what developer tool vendors should—and should not—do right now.

From Delivery Forms to the Selection Layer: No Matter How Good the Feature Promise, You Must First Cross the Admission Threshold

Over the past two years, software delivery forms have steadily shifted toward AI. In early 2025, code repositories began adding model-facing instruction files alongside human documentation, akin to onboarding guides prepared for new interns. Platforms like 21st.dev exemplify this pattern: developers input natural language, and the platform delivers UI component code directly. While traditional software libraries deliver building materials, AI-oriented component libraries deliver full-service construction crews. A previous analysis of Claude Code documented this delivery model.

By late 2025, vendors took this approach a step further: rather than bundling only functional code, they began concurrently delivering streamlined core suites, machine-readable assembly manuals, and companion tooling designed to reduce error rates. This entire paradigm is referred to as the generative kernel. Much like IKEA providing standard flat packs alongside the requisite Allen wrench, the generative kernel aims to deliver self-contained interfaces to automated tools. Its design principle is radical transparency: machine-facing design must expose raw errors and underlying parameters as much as possible, minimizing guesswork during model inference. When an API call fails, return the raw error fields and triggering parameters verbatim instead of swallowing them behind a friendly, generic message. A previous technical reflection on going beyond DRY explored this mechanism in depth.

Even if a vendor builds a machine-adapted generative kernel, a prerequisite question remains: what determines whether that kernel enters the agent’s field of view in the first place? Before a tool is actually assembled into a project, there must be a distribution interface that determines who gets into the engineering repository—and that is the selection layer. While product form determines whether a tool is pleasant to use after being chosen, the selection layer determines whether it can cross that first threshold into the engineering environment at all.

A high-stakes contest has already unfolded over who crosses this threshold. On the commercial services side, Armature launched optimization consulting starting at $5,000 per month to help developer tools increase their adoption probability in agent environments. On the open-source side, the project preseason.ai established a publicly adversarial benchmark dashboard by periodically running fixed prompt and model snapshot matrices.

Placing the test data from both sides side by side reveals a stark contrast. In the database category, preseason.ai testing showed PostgreSQL at 53.8%, Supabase at 24.9%, and Neon at only 5.3%, sharply divergent from Neon’s 66.3% commanding lead in Armature’s evaluation. As preseason.ai noted in its repository documentation, evaluation methodologies should be open, reproducible, and open to scrutiny, rather than locked away in private dashboards.

The methodological differences between the two test suites point to a plausible explanation: preseason asks models to respond directly, where answers largely track brand familiarity acquired during training; Armature has agents actually run integrations in a sandbox, where outcomes depend on documentation quality and the frictionlessness of the integration path. While this hypothesis stems from contrasting the two test designs and has not yet been validated via controlled comparison, it demonstrates that the competitive center of gravity for developer tools has expanded from raw code capability to agent-facing discovery and adoption.

There Is No Universal Leaderboard for Agent Tool Selection: It Is Fundamentally a Selection Function Conditioned on Task Context

The way coding agents acquire external information directly dictates the discovery paths for tools. In the test sessions, agents relied primarily on three channels. First is web search. In Armature’s reported statistics, the GPT-5.6-based Codex exhibited a web search invocation rate as high as 94%; in a sample of 57 searches, we found that 56 utilized the domain-limiting site: syntax. Second is built-in vendor knowledge skills. Some agents encapsulate dedicated documentation retrieval tools that bypass public search engines to read interfaces directly. Third is the repository’s local memory and pre-installed files. Agents prioritize adhering to the project’s existing lockfiles and configuration structures.

In contrast, Claude Code had an overall search invocation rate of about 30%, which only rose to around 80% in niche technical tests such as sandboxes. The industry offers two explanations for this. One view holds that more advanced models tend to rely on pre-trained general knowledge, thereby reducing external retrieval; another view emerges from community discussions, where active Hacker News engineer 42piratas pointed out that the discrepancy stems from permission friction within the framework itself: Claude Code must request user permission on a per-domain basis to fetch web pages and has an independent gate for search, making retrieval of existing repository files the path of least resistance, whereas Codex has no such friction. Multiple engineers reported that unless explicitly instructed to conduct research, Claude Code rarely searches the web proactively. Neither explanation is currently conclusive.

As for how machines actually read technical documentation, academic research and engineering deployments have already made the trends abundantly clear. An empirical study on AI agent documentation access behavior noted that multiple coding agents exhibit unique HTTP signatures when hitting documentation endpoints, compressing multi-page browsing into just one or two requests, which renders traditional metrics like time on page and bounce rate obsolete. Snowflake announced in release notes the launch of tiered llms.txt and Markdown documentation, explicitly designated for direct consumption by Cortex Code, Cursor, and Claude Code. A guide published by GitBook similarly supports returning structured Q&As directly via query parameters, noting that model selection systematically compares concept and constraint matrices.

Taken together, these touchpoints constitute an engineering-environment-driven selection function. An agent’s selection decisions are highly contingent on the specific conditions of a concrete task. Given identical email requirements, repositories across four programming languages converged on different providers; in voice interaction testing, simply toggling the persona to enterprise engineer shifted the preferred tool to LiveKit (23/76). Fine-tuning the prompt altered outcomes just as readily: introducing cost considerations into the prompt allowed Render to notch 30/54 victories, surpassing Vercel, which managed only 7/54 under the same conditions. On official model comparison dashboards, model version iterations similarly reshuffled selection rankings.

While the landscape in large markets remains relatively stable, long-tail niche markets are completely in flux. The preceding empirical data illustrates the resilience of market leaders: Stripe, with 349/395, and Neon, with 236/356, maintained solid leads across all three agents, buoyed by lucid interface specifications and minimal integration friction.

Yet in more fragmented long-tail segments, selections across different agents exhibited pronounced dispersion. The voice interaction benchmark mentioned earlier produced three distinct winners across the three agents: Claude Code favored Twilio (23/94), Codex favored OpenAI Realtime (32/110), and Cursor favored Vapi (41/104). LangChain, which appeared frequently across multiple evaluations with 197 official mentions, was ultimately adopted and integrated into the codebase only 4 times.

The Armature report also features another figure: a 42% selection agreement rate across tests. While the official definition of this metric was not disclosed, a reasonable deduction is that it represents the proportion of runs where all three agents chose the same tool for an identical task (same domain, same repository, same persona). Recalculating across the initial release dataset under multiple sensible definitions yielded a range between 41.5% and 48.2%, numerically consistent with Armature’s published 42%. This distribution indicates that vendors cannot expect to lock in static model preferences; what can be sustainably managed is only reachability within specific task contexts.

SEO optimizes for global rankings, whereas agents select winners conditioned on task contexts

Analogizing Agent Selection to SEO: Which Lessons Apply, and Which Logic Breaks Down Completely

Many instinctively analogize agent-oriented tool optimization to traditional search engine optimization (SEO). At a macro level, the analogy holds water: software procurement has introduced new intermediaries and distribution channels, spawning corresponding optimization services and measurement markets. In a community discussion on Hacker News, the study garnered 300 points and 152 comments, with many developers voicing concern over a proliferation of new spam. Co-founder Louis Scremin replied that this could escalate into a mess a thousand times worse than traditional SEO, urging the industry to tread carefully.

When dissecting the concrete engineering mechanics, however, this analogy breaks down across three distinct dimensions. First, the mathematical formulation of the optimization target is fundamentally different. Traditional search centers on a globally stable ranking function that generalizes across users, allowing vendors to optimize against fixed ranks over long horizons. In contrast, a coding agent’s decision hinges on specific task contexts—it is inherently a selection function. The same tool that wins in a TypeScript repository may fail in a Python repository; injecting cost constraints into a prompt can prompt an agent to pivot. Vendors cannot optimize for a static global rank; the goal shifts to satisfying execution requirements within concrete engineering contexts. This underpins Armature’s business model of selling benchmarks segmented by repository, language, and persona.

Second, safety guardrails establish a technical ceiling. In the traditional search era, websites could manipulate rankings through keyword stuffing. In agent interactions, heavy-handed steering directly triggers security defenses. Armature co-founder Louis Scremin revealed an unpublished experiment: when the team applied excessive bias toward a specific product in an internal search engine, it immediately triggered prompt-injection defenses. Frontier models would rather incur false positives than overlook risks during safety evaluations. While objective, factual information supplies actionable decision grounds, forced manipulation triggers defenses and results in direct interception.

Third, the accountability chain for procurement decisions has not yet been handed over. The aforementioned figure of over 50% of deployments is strictly an execution-layer metric; account credentials and payment authority remain firmly in the hands of engineers. While the simulated user in testing ultimately approved all proposals, it disregarded existing enterprise vendor agreements, security compliance mandates, and legacy stack constraints. In production environments, humans remain the final filtration layer for software selection.

Another point often invites confusion. Many assume that adding an llms.txt file to a website guarantees citations by large language models, but two independent datasets refute this claim: Google’s official guidelines explicitly state that search does not utilize such formats, and Search Engine Journal’s statistical study of approximately 300,000 domains identified no positive effect on search citations. Furthermore, bibliographic citation in general search engines and code integration by coding agents follow two distinct pathways; one should neither wager on the latter’s maturity based on the former, nor use the former’s immaturity as an excuse to ignore the latter.

Deliverable forms evolve layer by layer, with competition emerging at the outermost selection layer

Conclusion

In response to the emergence of the selection layer, developer tool vendors need to establish a fresh measurement perspective. The tool conversion funnel can be decoupled into four distinct stages: first, whether an agent can locate and read about the tool during retrieval; second, whether the tool enters the candidate shortlist; third, whether the agent can successfully write code to integrate it in a specific repository; fourth, whether the integrated code passes tests and satisfies functional requirements. Armature’s benchmarks conflate discovery and integration friction; vendors must measure these four stages independently when conducting evaluations.

Synthesizing inferences from current evaluations, vendors can advance four tactical steps in documentation development. These recommendations represent extrapolations from the current landscape rather than ready-made panaceas:

First, clearly delineate applicability boundaries. Explicitly state supported framework versions and unsuitable scenarios in documentation to prevent agents from attempting blind integrations.

Second, disclose realistic pricing and quota limits. Test logs indicate that agents scrape and comprehend pricing structures: Mailgun lost to Postmark because its free tier noted a one-day log retention window, and Supabase fell short in standalone database evaluations because its pricing bundled auxiliary backend services.

Third, provide verified, minimal runnable examples. Supply zero-configuration snippets to minimize dependency conflicts.

Fourth, reduce integration friction. Minimize interactive configuration steps and permission gates, enabling code to build silently inside sandboxes.

When refining documentation, vendors must delineate strict security boundaries. Web content should be confined to providing objective facts without crossing prompt-injection red lines. Because distribution entry points represent potential attack surfaces, documentation may articulate technical constraints, but must never attempt to solicit execution privileges.

Vendors must also exercise caution regarding the efficacy of copy optimization. In the deployment benchmark, all 18 test cases altered their selections based on changes in prompt phrasing, measuring responses triggered by variations in user phrasing. The original experiment was designed around altering prompt phrasing under identical repository and agent conditions, rather than a controlled A/B test of vendors updating their landing page copy. Rigorous empirical data is currently lacking on the extent to which on-page copy adjustments causally alter adoption rates.

Before progressing toward truly automated procurement, the industry confronts three unresolved questions. First, in real-world engineering, to what extent do agents possess autonomous selection authority versus merely executing installations designated by humans? Second, how much statistically significant causal lift can vendor modifications to documentation copy actually deliver to adoption rates? Third, as agentic capabilities expand, will engineers gradually forgo line-by-line review in low-risk domains and genuinely cede procurement authority to automated workflows?

In an earlier discussion on software engineering, a perspective was raised: software engineering is shifting from building concrete software to engineering its generative potential. Today’s reality deepens that realization a step further: the generative potential latent within software must first be discoverable, intelligible, and seamlessly executable by agents within a concrete task context. This discovery mechanism linking latent capability to engineering practice constitutes the entirely new selection layer in software distribution.