Cold Calling Data Providers and Dialers That Publish Outbound Benchmark Research
Vendors publishing cold calling benchmarks also sell the products their research validates.

Cognism, Gong, Belkins, and Orum each publish research on cold calling performance, and each also sells a product that the research tends to validate. That is the subject of this piece: the companies with the largest call datasets in the industry are the same companies selling into the market those datasets describe. A benchmark showing strong connect rates for verified mobile numbers tends to come from a company selling verified mobile numbers. A benchmark showing that parallel dialers multiply live conversations tends to come from a company selling parallel dialers. None of this makes the underlying study false. It does mean the sample a publisher chooses, the metric it highlights, and the denominator it builds that metric on are decisions shaped by what the publisher needs the number to prove. A sales team that takes one of these figures and writes it into a quota plan is, often without realizing it, accepting someone else's definition of what counts as a connect, a success, or a meeting. The sections that follow work through who these publishers are, what their samples contain, how their definitions diverge, and what gets left out of the connect-rate conversation.
Who the primary benchmark publishers are
A small group of organizations produces most of the cold calling research that circulates through sales blogs, LinkedIn posts, and planning decks, and each one sells a product tied to the category it studies.
Belkins published "B2B Cold Calling Benchmarks 2026," built from more than 175,000 dials logged throughout 2025. The study reports both per-dial and per-prospect connect rates. Belkins runs outbound campaigns for other companies as an agency, so its business depends on showing that outsourced SDR work turns into measurable meetings.
The Bridge Group's 2025 SDR Metrics report draws on 351 B2B companies and publishes daily activity benchmarks across phone calls, emails, and LinkedIn touches. The Bridge Group sells SDR consulting services alongside its benchmarking research.
Saleshandy built a dataset of 7,699 calls collected between July and August 2026, drawn entirely from users of the Saleshandy Dialer. Cognism, a provider of verified contact and mobile data, publishes its own State of Outbound research built from the calling behavior of teams already using its platform.
A separate category comes from independent practitioners with no product riding on the result. The Outbound Kitchen funded its own 2026 mobile data benchmark, testing several providers against the same 1,400 contacts, with no provider paying for inclusion. That funding structure matters for how the conclusions should be read, and the next three sections work through what each kind of study can and cannot tell a reader.
How each publisher's sample shapes its benchmark
A benchmark is only as wide as the population it was drawn from, and the population behind most cold calling studies is the publisher's own customer base.
Cognism's dataset reflects teams already using Cognism's verified mobile data, not teams working unverified or generic contact lists. Its headline success-rate figure describes what happens for Cognism customers specifically. A team buying numbers from a different source should not expect to land in the same range. Belkins' dataset comes from its own agency's campaigns, so it captures the performance of outsourced outbound run by specialists who do this calling for a living, not an internal SDR team at a five-person startup or a founder dialing a list between meetings.
Saleshandy's 7,699 calls came entirely from people who had already chosen to use the Saleshandy Dialer, a self-selected group of active outbound practitioners rather than a random cross-section of cold callers generally. The Bridge Group's 351 companies self-reported their own activity levels to researchers. The Bridge Group itself acknowledges these are anonymous operator reports, not call logs pulled from a system. None of these samples is biased in the sense of being rigged. Each one is bounded: it describes a specific population accurately, and the open question for any reader is whether that population resembles their own team, their own list, and their own market.
Sample size gets the most attention in how these studies get marketed, but sample composition does the real work of determining whether a number applies. A dataset of 175,000 dials from an outbound agency and a dataset of 7,699 calls from a dialer's user base are not comparable in scale, but scale is not what decides relevance here. A large dataset built from a population unlike the reader's own team tells the reader less than a small dataset built from a population that matches it closely, and the question worth asking is which study's callers most resemble a given team calling into a given market, with a similar list, a similar product, and a similar level of calling experience. The LevelUp Leads benchmark guide states this directly: neither figure from its compiled benchmarks should be copied into a forecast without first matching the denominator and the campaign context to the reader's own situation, a disclosure that is rare in vendor-published research.
Why the same metric means different things across studies
Two honest studies can report the same metric and produce numbers that appear to contradict each other, because the metric rests on a different denominator in each case.
Saleshandy's own benchmark guide identifies five separate metrics in circulation: success rate, connect rate, contact rate, meeting conversion rate, and conversation success rate. The guide notes that these terms get used interchangeably across the industry, which makes cross-study comparison unreliable even before any numbers enter the conversation. Belkins' study demonstrates the problem concretely. From the same dataset of logged dials, Belkins reported a 9.9% per-dial connect rate alongside a 24.5% per-prospect connect rate. Belkins describes the gap between the two as essentially the value of a follow-up call: the per-prospect figure counts each prospect once, no matter how many attempts it took to reach them, while the per-dial figure counts every single attempt separately. A team that takes the 24.5% figure and compares it against another study's per-dial rate is measuring two different things and calling them the same metric.
The LevelUp Leads compilation adds a second layer to the Belkins numbers: Belkins counted any live human answer as a connect, including gatekeepers, wrong-number pickups, and immediate rejections, while counting something as a conversation only once the contact stayed on the line past the opening line. Connect rate and conversation rate, inside the same study, describe two different outcomes. Treating either figure as a stand-in for the other produces a target nobody can actually hit.
The same fragmentation separates Cognism from Saleshandy. Cognism's success-rate denominator is total calls made. Saleshandy's connect-rate denominator is total dials placed to verified numbers only. A team working unverified data that uses Cognism's number as a floor and Saleshandy's number as a ceiling is, without meaning to, measuring itself against two separate populations of dialed numbers. None of this reflects dishonesty on the part of any publisher. It reflects definitional fragmentation across an industry that has never agreed on what these words mean, and that fragmentation gets amplified the moment a headline number leaves its original report. A figure that enters a blog post, a sales deck, or a planning spreadsheet almost always loses its denominator along the way. LevelUp Leads states the safer practice: carry the source, the sample, and the definition alongside every number, treating each figure as a single reference observation. The Belkins example shows why that discipline matters. The same 175,000 dials produced a connect rate nearly two and a half times higher once the denominator switched from dials to prospects, and both numbers were accurate.
What the data quality studies reveal
Most connect-rate benchmarks treat the quality of the phone list as a fixed, unstated assumption. Independent research into data quality shows that assumption is often where the real variation lives, ahead of rep skill, script, or dialer choice.
The Outbound Kitchen's 2026 benchmark ran the same 1,400 US contacts through several data providers at once. The widest provider returned a valid mobile number for the large majority of those contacts. The most accurate provider was correct only about two-thirds of the time. No single provider sat in the top-right corner of both measures simultaneously, meaning wide coverage and high accuracy did not come from the same source. Across every provider combined, one in four valid-looking mobile numbers actually belonged to the wrong person. Of every number confirmed valid, roughly three in five reached the right person, about a quarter reached someone else entirely, and the remainder could not be confirmed without placing the call.
That gap changes what a connect-rate benchmark can promise a reader. A team dialing a list built from a single provider, then applying a connect-rate target drawn from a study that ran on verified data, is planning against a number that assumed cleaner inputs than the team actually has in hand. Salesfinity's controlled benchmark across nine providers illustrates the trade-off from the dialer side of the business: high accuracy without wide coverage limits how many people a team can reach at all, while wide coverage without accuracy inflates the number of calls wasted on wrong numbers. Salesfinity sells a dialer with its own enrichment features built in, so its benchmark is best read as an illustration of this trade-off rather than a neutral audit of the market, and it reaches a similar conclusion to the independent Outbound Kitchen research: no single vendor wins on both dimensions, only specific combinations suit specific segments.
The practical effect is that most published connect-rate benchmarks are not benchmarks for cold calling in general. They are benchmarks for cold calling performed against one particular tier of data quality, and that tier is rarely stated next to the headline number. Cognism's own finding, that verified contact data paired with AI-assisted prioritization pushed cold call answer rates to 13.3%, nearly matching the 14.4% answer rate AEs get calling warm leads, is a real result for Cognism-verified data specifically. It is not a result for phone data as a category. A team dialing a list from a provider that returns a high share of wrong-person numbers will not reach that 13.3% mark no matter how well-trained its reps are or how capable its dialer is, because the ceiling was set by the list before the first call was placed. The Outbound Kitchen study puts an actual cost on that gap: bad data costs a ten-rep team a substantial sum each year in wasted rep time alone, before any accounting for the pipeline that call never had the chance to build.
Dialer vendors and product value in published metrics
Dialer vendors tend to publish metrics that measure the output of dialing activity itself: conversations per hour, talk time gained, dials completed per rep per day. These are the numbers a dialer can move directly, which is also why they're the numbers dialers choose to report.
Retell AI's 2026 comparison of the dialer market evaluates providers across dialing mode, covering parallel, power, predictive, and AI agent-plus-batch approaches, alongside voice quality, latency, compliance controls, scale economics, and ease of setup. Every one of those axes describes something a dialer's engineering and configuration can change directly. None of them describes what happens after the call connects, whether the prospect was the right person, or whether the conversation turned into pipeline.
Saleshandy's own dataset illustrates the gap. Saleshandy Dialer users calling pre-verified numbers reached a notably high connect rate in the company's study. That figure is accurate for exactly the population it measured: existing Saleshandy users working lists that had already been verified. A new user should not expect to hit that figure before their own list has gone through the same verification step, a gap that reflects differing populations rather than any vendor inventing numbers or hiding what it measures. Talk time and conversations-per-hour are genuine intermediate inputs. A rep who spends more time in live conversation has more chances to convert that conversation into a meeting, and a dialer that increases live-conversation volume is doing something real, because activity metrics and outcome metrics answer different questions, and a dialer is transparent about which one it measures.
The structural limit on any dialer's published pickup rate is that the dialer cannot control for the data quality each customer brings to it. A pickup rate published by a dialer vendor reflects the average data quality across that vendor's own customer base at the moment the study was run. A team adopting a faster dialer without first addressing the accuracy of its contact list is optimizing one variable in a chain where the list, not the dialing speed, may be setting the real ceiling on what connects are possible.

