Performance marketing · AI search

You can now buy ChatGPT ads through Amazon DSP. Should you take that route?

There are now two ways to put an ad inside ChatGPT, and as of last week they are not equivalent. One is a self-serve platform with a conversion pixel and conversion bidding. The other is a managed extension of your Amazon DSP campaigns that reports delivery metrics and not much else. Which one you should use depends less on which brand you trust and more on what you intend to prove.

The short answer

Amazon Ads announced on 10 September 2026 that select US advertisers can extend Amazon DSP campaigns into ChatGPT Ads, managed by Amazon reps, with aggregated reporting on impressions, clicks, CPM, CPC and cost per result. Buying direct through OpenAI gives you conversion-optimised bidding, a conversion pixel and mobile measurement integrations. Choose Amazon for audience reach and one workflow; choose direct if the test has to prove conversions.

What Amazon announced

On 10 September 2026, Amazon Ads announced an integration with OpenAI that lets advertisers extend their Amazon Ads campaigns into ChatGPT with a conversational ad format. The specifics that matter for a media decision:

The strategic logic is straightforward on both sides. Amazon gets to sell inventory it does not own to demand it already has, continuing the pattern it set by merging DSP and Sponsored Ads into a single account in July 2026 (covered in what that merger changed). OpenAI gets access to Amazon’s advertiser base and its commerce signals without building those relationships itself.

The context for both: OpenAI’s advertising business passed a $1 billion annualised run rate at the end of August 2026, less than 200 days after launch, against more than a billion weekly ChatGPT users of whom roughly 20% show commercial intent. This is no longer an experiment you can dismiss on scale grounds. It is an experiment you should evaluate on measurement grounds.

Three routes, compared

There are now three defensible positions, and the honest answer for most advertisers is still the third one.

Amazon DSP managed pilotDirect through OpenAIWait
AccessSelect US advertisers, invitation-basedSelf-serve, 40+ countries—
Who operates itAmazon Ads reps, managedYour team or your agency—
Targeting controlIndirect, via context hints supplied with Amazon’s helpDirect, in the platform—
BiddingDelivery-led; CPM and CPC reportedConversion-optimised CPC available since July 2026—
Conversion trackingNot part of the reported pilot metricsConversion pixel plus mobile measurement partner integrations—
Workflow costLow. It sits inside a buy you already runHigher. A new platform, new creative, new pacingZero
Best suited toBrands already spending in Amazon DSP who want incremental reachAdvertisers who need to prove a conversion outcomeAnyone whose search and social are not yet at diminishing returns

Note what the table does not say. It does not say the Amazon route is worse. For a brand running a $400,000 quarterly DSP buy, adding a new surface through an existing rep relationship at no extra operational cost is a reasonable thing to do, and the reporting is adequate for the job that buy is doing. It says the two routes answer different questions.

The measurement gap is the decision

Almost every write-up of this announcement led with access. Access is the least interesting part, because the inventory is the same inventory. What differs is what you can know afterwards.

What you want to knowAmazon DSP pilotDirect through OpenAI
Did it deliver?Yes — impressions, CPMYes
Did anyone engage?Yes — clicks, CPCYes
Did it produce a defined result, at what cost?Cost per result, for whatever result the pilot countsYes, against your own conversion definition via the pixel
Can the platform optimise toward your conversion?Not in the reported pilotYes — conversion-optimised CPC
Does app install or in-app revenue flow back?Not statedVia mobile measurement partner integrations
Independent verification and viewabilityNoNot yet; OpenAI has said it is coming without naming partners or dates

Two traps in that table deserve naming.

“Cost per result” is not a CPA. A result is whatever the platform counts as one, and on a conversational surface that could be a click-through, a completed interaction or an assisted action. Before that number enters a spreadsheet next to your Google Ads CPA, ask the rep to define the event in one sentence. If the definitions differ, the comparison is not a comparison, and this is precisely how new channels acquire undeserved reputations in both directions.

Aggregated reporting caps what the test can conclude. With delivery metrics only, the strongest honest claim is “we reached this audience at this cost”. That is a legitimate claim. It is not “this channel drove revenue”, and a quarterly review that quietly upgrades the first into the second is how budget gets misallocated for a year.

How to read a test when attribution will not do it for you

If you take the Amazon route, platform attribution is not going to answer the business question, so design the test to answer it without attribution. A geo holdout is the practical tool, and the discipline is the same as in any incrementality test: decide in advance what size of effect you could detect.

A worked example, with the arithmetic that most test plans skip:

  1. Pick the paired geographies. Twenty comparable metros, split ten and ten, matched on the last 13 weeks of revenue rather than on population.
  2. Establish the noise floor. Take weekly revenue in your would-be test geos for the last 13 weeks and compute the week-to-week variation. Suppose baseline weekly revenue is $250,000 and normal variation is about ±4%, or ±$10,000.
  3. Work out what you could see. Over a six-week test, random variation of that size means effects smaller than roughly 2–3% of revenue are indistinguishable from noise. On a $250,000 base, that is about $5,000–$7,500 a week, or $30,000–$45,000 over the test.
  4. Compare that with the spend. If the test budget is $40,000, you are asking whether $40,000 of spend produced at least $30,000–$45,000 of incremental revenue — roughly break-even at best. A test that can only detect an effect larger than the one you would be happy with is not a test, it is an expensive coin toss.
  5. Fix it one of three ways: concentrate the same budget in fewer geos to raise spend per market, run for ten weeks instead of six, or accept up front that the read is directional and label it that way in the deck.
What we’d do about it

Whichever route you pick, add one measurement that costs nothing: branded search volume and direct traffic in test versus control geos, weekly. Conversational surfaces tend to create demand that lands somewhere else, so an ad inside ChatGPT can show up as a branded search three days later. If you only watch last-click, you will conclude the channel did nothing.

Choose X when, choose Y when

The third bullet applies to more businesses than the first two combined, and there is no prize for being early to inventory that will still be there in January — larger, cheaper to measure and with more competitor data to learn from.

The Q4 timing problem

This landed in the worst possible week of the year to add a variable. Black Friday is 27 November 2026 and Cyber Monday is 30 November. Between now and then, Google is auto-migrating Search campaigns to AI Max, standalone Display campaigns are moving into Demand Gen, and language targeting for Search and Performance Max changes from late September.

A new channel on a managed pilot with unfamiliar reporting is a poor fit for the six weeks when your existing channels are already being rewritten underneath you. If the pilot invitation is open, take the meeting, get the “result” definition in writing, build the landing experience, and run the test in January when the baseline is stable and a 3% lift is not hidden inside seasonal variance of 40%.

What this signals beyond one pilot

The pattern worth noticing is that AI answer inventory is being sold through existing demand-side relationships rather than only direct. Amazon selling ChatGPT placements is the same structural move as retail media networks selling off-site inventory: the buyer stays where their budget and their rep already live, and the new surface arrives as a line item.

For advertisers, that has one practical consequence. You will increasingly be offered AI-surface inventory inside buys you already run, described in the reporting language of the platform selling it rather than of the platform serving it. The defence is not scepticism about the channel; it is a habit of asking three questions before each new line item: what event is being counted, who counts it, and what would I have to see to keep spending here next quarter. That is the same discipline that performance marketing has always needed, applied to inventory that is currently better at selling itself than at proving itself.

Frequently asked questions

What exactly did Amazon announce on 10 September 2026?

Amazon Ads announced an integration with OpenAI that lets a select group of US advertisers extend their Amazon Ads campaigns into ChatGPT with a conversational ad experience, bought through Amazon DSP as a managed service. Amazon Ads staff help build and manage the campaigns and supply what Amazon calls context hints, which match ads to relevant topics, conversations and keywords. Delta Vacations is a named pilot advertiser.

Is ChatGPT Ads inventory the same on both routes?

The inventory is the same surface, but the buying experience and reporting are not. The Amazon route is a managed pilot limited to the US with aggregated delivery reporting. OpenAI's own platform, which added conversion-optimised CPC bidding, a conversion pixel and mobile measurement partner integrations in July 2026, is available in more than 40 countries and lets you optimise toward a conversion objective yourself.

What can you actually measure on the Amazon route?

According to the pilot terms as reported, aggregated performance data covering impressions, clicks, cost per result, CPM and CPC. That is delivery and cost reporting, not outcome measurement. Cost per result is not a cost per acquisition unless the result being counted is your acquisition, so establish what event that metric refers to before you use it in any comparison.

How big is ChatGPT Ads now?

OpenAI's advertising business passed a $1 billion annualised revenue run rate at the end of August 2026, reported by CNBC on 31 August 2026, which is roughly $83 million a month and less than 200 days after launch. ChatGPT has more than one billion weekly users, of whom OpenAI says about 20% show commercial intent. It is real inventory at meaningful scale, and it is still young enough that measurement standards are unsettled.

Should a small business test ChatGPT ads at all in Q4 2026?

Only with money it can afford not to attribute, and only if search and social are already at diminishing returns. A defensible first test needs a budget that can produce a readable result, a landing experience built for conversational traffic, and a measurement plan that does not depend on platform attribution. If any of those three is missing, the test will produce an anecdote.

Does third-party verification exist for ChatGPT ad inventory?

Not yet in the form advertisers are used to. OpenAI has said third-party measurement is coming but has not named partners or a timeline, and the viewability and verification ecosystem that surrounds Google and Meta inventory does not exist here. For brands with verification requirements in their media standards, that is a reason to wait rather than a detail to negotiate.

The takeaway

Two routes into the same inventory, with different jobs. The Amazon DSP pilot is an audience-extension buy for advertisers already spending in Amazon DSP, priced and reported like an awareness placement, with a rep doing the work. Buying direct from OpenAI is the performance route because it is the only one that currently lets you optimise toward a conversion and see it. Pick the route that matches the claim you need to make at the end of the quarter, size the test so it can actually support that claim, and if neither condition holds, spend the money where the measurement already works and revisit in January.

Rahul Gupta

Founder of HyberX, a digital growth agency working with brands across the US, Europe, the Middle East and India. Writes on web design, paid media and conversion optimisation.

More about Rahul · LinkedIn

Related reading

Want a ChatGPT ads test that produces a decision, not an anecdote?

We size the test, pick the buying route, build the holdout so the result is readable, and tell you plainly when the answer is to wait.

Book a Growth Call