You've successfully subscribed to Circleboom Twitter: Analytics & Management for X Accounts
Great! Next, complete checkout for full access to Circleboom Twitter: Analytics & Management for X Accounts
Welcome back! You've successfully signed in.
Success! Your account is fully activated, you now have access to all content.
Twitter ad agency vs in-house: the decision I watched go wrong twice

Twitter ad agency vs in-house: the decision I watched go wrong twice

. 8 min read

A peer of mine spent most of a year cycling between two Twitter ad agencies, and the second one lost money faster than the first. Not because it was worse. Because it was better at spending, and nobody had fixed the thing underneath.

The thing underneath was that both agencies were buying the same audience, from the same interest categories, at the same moment in the buyer's decision. Changing the vendor changed nothing about who saw the ad.

I watched that happen, decided I was too smart for it, and then made the mirror-image mistake myself by pulling everything in-house and losing the only part of the arrangement that had been working.

Both errors came from the same wrong question. I kept asking who should run the campaign.

The question that mattered was which layer of the campaign was actually broken.

What you are really deciding between.Agency judgment on message, creative, and testing cadence, which is genuinely hard to replace.Audience construction, which is a keyword search, a bot filter, and a CSV export.Exclusion work, which almost nobody does at all, and which is where the fastest savings sit.

Circleboom handles the second and third of those on X through official Enterprise access. Start with the real-time tweet tracker to see who is discussing your category right now.

The stakes nobody puts in the proposal

Here is the arithmetic that should sit on the first page of every Twitter ad agency proposal and never does.

Take a modest monthly media budget of $2,000 and a retainer of $1,500 on top. If broad targeting delivers impressions where roughly one in ten engagements comes from a real, relevant human, then the effective media spend on reachable people is around $200 a month.

You are paying $1,500 in fees to manage $200 of useful reach.

That is not an argument against agencies. It is an argument against evaluating one on media efficiency when the audience layer is unfixed.

Fix the denominator first, then decide who manages the numerator.

Global ad spending forecasts keep climbing, with social capturing a growing share of digital budgets according to eMarketer's worldwide ad spending forecast.

More money flowing into the same broad-targeting defaults does not improve the ratio. It just makes the leak larger.

Before any renewal conversation, read the line-by-line version. The Twitter ads agency quote that came to $3,000 a month shows which parts of a retainer survive that arithmetic and which do not.

Before you can price any of it, though, you need to know what the work produces. Watching your own category through real-time tweet tracking for a week gives you that baseline, and it costs nothing but the week.

The layer nobody sells you: exclusion

Every agency proposal I have read talks about who to reach. Not one has talked about who to stop reaching, and that omission is expensive in a way that compounds silently.

X supports a negative-targeting surface most advertisers never touch. Its Do Not Reach Lists documentation describes uploading a list of accounts that X excludes from delivery entirely, independent of whatever positive targeting you have set.

Think about what belongs on that list:

  • Existing customers who are already paying you.
  • Your own employees and their alternate accounts.
  • Competitors and industry press who monitor your category.
  • Accounts that engaged, converted, and are now in a support relationship.
  • The bot and low-quality profiles you identified while building your positive list.

That last one is the interesting one, because building the positive list produces the negative list as a by-product.

When you filter a keyword search down from 5,000 raw accounts to 1,200 usable ones, the 3,800 you rejected are not garbage. They are an exclusion file you already paid to generate.

Most advertisers throw away the more valuable half of their own filtering work.

I did this for a year. I built careful positive lists, uploaded them, and let the rejected accounts evaporate.

Meanwhile the campaign was still free to serve to every one of them through any broad-match component in the setup.

The economics of that are worth sitting with. Filtering is the expensive part of audience work, because it is where judgment gets applied and where the data access earns its keep.

Discarding the rejected half means paying for the expensive step twice: once to identify who not to reach, and again in media spend reaching them anyway.

How I run a Twitter ad agency evaluation now

The process I use before signing, renewing, or bringing anything in-house. It takes an afternoon and it decides the question for you.

Build the intent list from live conversation

  1. Connect your X account inside Circleboom Twitter. This is the only step that touches your credentials, and it happens once.
  1. Go to Advanced X Search and choose the real-time tracker rather than the account search. These sit next to each other and do genuinely different jobs: the account search reads bios, the tracker reads what people are posting right now.
  1. Track the phrases buyers use when they have the problem, not the phrases your category uses to describe the solution. Real-time tweet tracking shows you accounts discussing a topic as it happens, which is a fundamentally different signal from a static bio keyword.
  2. Collect the accounts behind those posts and let the pool build for a week before you touch it. Intent expressed once is noise; intent expressed repeatedly is a segment.

Split the pool into a target list and an exclusion list

  1. Apply the quality filters to the collected accounts: exclude Fake/Spam, Inactive, Eggheads, and Protected, then set follower-count and follow-ratio bounds that match a plausible buyer.
  2. Export both halves as CSV. The accounts that passed become your Custom Audience; the accounts that failed become your Do Not Reach list. Two files from one filtering pass.
  3. Upload both in X Ads Manager and run the campaign against the positive list with the negative list applied.

Doing it in this sequence is what stops the evaluation from being a negotiation with yourself.

You now know how long the audience work takes, what it produces, and what it costs, which means you can price the agency's version of the same job instead of guessing at it.

An agency that adds real value on top of that file will be obvious within one campaign cycle, and so will one that does not.

See it live: the real-time keyword tracking that turns an active X conversation into an audience file.

What changed after I fixed the layer instead of the vendor

The first campaign I ran with both lists applied behaved differently in a way I did not expect:

  • Impressions dropped noticeably.
  • Frequency on the remaining audience rose.
  • Cost per click got slightly worse.

By every dashboard metric, it looked like a downgrade.

The conversations coming out of it were not a downgrade. Replies came from accounts that had posted about the problem in the previous fortnight, and several of them referenced their own post when they responded.

Intent-based targeting produces worse-looking metrics and better-looking outcomes, and you have to decide in advance which one you are managing to.

That decision has to be made before the campaign runs, because once the numbers are in front of a room, the worse-looking metric always wins the argument.

Here is what I got wrong. I treated "I can produce the file" as evidence for "I can run the programme," and those are unrelated claims.

The file took an afternoon. The thing the file feeds, a rotation of creative tested against itself week after week, is a standing commitment, and I had budgeted no standing commitment at all.

I had diagnosed a vendor problem and bought a scope problem instead.

Auditing lists is also where two axes I had been ignoring turned out to matter. If your instinct right now is that X ads are simply expensive, why Twitter ads cost so much and where the cost actually goes is what reframed it for me.

The expense turned out to be a targeting problem wearing a pricing costume. That reframe is why I stopped negotiating rates and started auditing lists.

Location adds another axis worth using. Searching X by location narrows an intent pool to a market you actually sell into, which matters enormously if your product is regional and your keyword is not.

For ongoing monitoring rather than one-off pulls, the keyword and hashtag tracker keeps the intent pool refreshing on its own, so the list does not go stale between campaigns.

Four situations, four different answers

It depends on which layer is broken, and you can now diagnose that yourself.

If your targeting is broad and nobody has ever shown you an audience file: the agency is not your problem yet, and switching vendors will not help. Build the list and the exclusion list first, then re-run the evaluation. This was my peer's situation and it cost a year.

If your audience file is clean but your creative is stale and your testing has stopped: hire the agency. This is exactly the work that benefits from people who do it daily, and it is the work I proved I could not sustain alone.

If both are broken: fix the audience layer first anyway, because good creative against a bad list is unmeasurable. The new agency's work becomes untestable, and you spend the quarter unable to say whether it helped.

If both are healthy and you are simply spending more than you want: the answer is scope reduction, not vendor change. Take the audience line off the retainer, keep the creative and management line, and expect the total to fall without the output falling with it.

Small budgets change the calculus further. The practical constraints in running Twitter ads for small businesses make the agency question mostly moot below a certain spend, where the retainer alone would exceed the media.

Below that line the question stops being who manages the budget and starts being what to do with it. The tactics in creating effective Twitter ads on a small budget cover what actually works when the media budget is the constraint rather than the management.

And for the strategic frame underneath all of this, nanotargeting on X argues the behavioral case for why precision beats demographics on this platform specifically.

Circleboom sits in the audience and exclusion layers as an X Enterprise data partner, which means both files are built on sanctioned data rather than scraped profiles.

That distinction matters most on the exclusion side, where an incomplete list quietly fails to exclude the accounts you most wanted gone.

Whichever branch you land on, run the live tweet tracker once before you decide. The evaluation is free and the answer stops being a matter of opinion.

→ Start tracking real-time tweets in your category

Common questions about X ad agency work

Should I build the exclusion list before or after the target list?

Neither. Build them in the same pass. The accounts your filters reject while producing the target list are the exclusion list, so treating them as two separate projects doubles the work and usually means the second one never happens.

Will my agency object to me building the audience myself?

Watch what they ask about rather than whether they agree. The useful reaction is a question about your filter logic, because that is somebody thinking about whether your file is any good. The unhelpful reaction is a paragraph about methodology, which is almost always a way of not answering.

How long should I collect real-time tweets before exporting?

About a week for most categories. A single day of collection over-indexes on whatever happened to be trending, while a week smooths that out and lets repeat posters show up, which is the signal you actually want.

Does excluding accounts shrink my reach too much?

It shrinks reported reach, yes, and that is the point. The accounts on a well-built exclusion list were never going to convert, so removing them lowers impressions without lowering outcomes. Expect your dashboard to look worse before your pipeline looks better.

Can I run this evaluation without any ad budget at all?

Yes. Building the target list, the exclusion list, and the row counts costs nothing in media. You only need budget once you decide to run the campaign, and by then you already know what the agency's version of the audience work is worth.


Arif Akdogan
Arif Akdogan

Passionate digital marketer helping grow through innovative strategies, data-driven insights, and creative content. [email protected]