The short version: you can separate an operator from a salesperson in one call by asking ten specific questions and scoring the specificity of the answers. This guide gives you the ten questions, what a real answer sounds like for each, a simple scoring rubric, and the red and green flags around the process. We are an agency, so run the rubric on us too: that is what it is for, and every process we would be scored on is published on this site.
Before any calls: audit your own account
The single highest-leverage step in agency evaluation happens before you talk to anyone: run the 30-minute self-audit on your own Klaviyo account and write down your three biggest problems privately. Every agency you talk to will diagnose your account on the first call. The candidates worth hiring will independently find the problems on your private list; the ones to avoid will diagnose whatever their package sells. You cannot run that test if you have not looked first.
The 10 questions, and what a real answer sounds like
1. Which attribution window do you report on, and why?
An operator names a number, explains the trade-off, and volunteers that Klaviyo's defaults flatter the channel. A salesperson says "we report everything transparently" without a number. This question goes first because it predicts every future monthly report.
2. Walk me through your welcome flow structure for a brand like mine.
Listen for branching: buyers versus non-buyers, offer economics, timing logic. A real answer sounds like our published welcome build: specific splits, specific reasoning. A generic answer is "three to five emails introducing the brand."
3. What would you fix in my account first, and can you look now?
The strongest candidates will open your account (or a screen share) on the sales call and point at something real. Agencies that resist looking before proposing are pricing a package, not your problem.
4. How do you decide campaign cadence, and what would you set mine to?
A real answer asks about your engaged list size and season before naming a number, because cadence is a function of those inputs. Our version of the function is published with a calculator. A generalist answers with their standard package volume.
5. What does your deliverability monitoring look at weekly?
Listen for: spam complaint rate against the 0.1% ceiling, unsubscribe trend per send, click reach, and engagement-segment health. "We monitor deliverability closely" without named metrics means nobody is monitoring it.
6. How do you segment differently for my category?
Vertical mechanics are where category experience shows and cannot be improvised: replenishment timing for consumables, seasonal winback for outdoor, browse-first flows for considered purchases. The vertical questions in our agency guide and its vertical deep-dives give you the category-specific versions.
7. Who writes the copy and builds the designs, and how do they learn our voice?
You are evaluating whether the people in the pitch are the people on the account. Ask for the actual names, the onboarding process for brand guidelines, and revision turnaround.
8. What happened with the last client you lost?
Every agency loses clients. An honest answer with specifics is a green flag of the highest order; a claim of near-zero churn or a rehearsed non-answer tells you how they will communicate when your account struggles.
9. What do you need from my team monthly for this to work?
Real engagements need product information, approvals, and asset access on a rhythm. An agency that claims to need nothing from you is planning to run template output.
10. How will we know at day 90 whether this is working?
The answer should name metrics (flow revenue share, repeat-purchase movement, revenue per recipient) with attribution windows held constant, and it should not promise a transformation in month one, because architecture results take about 90 days of send volume to prove.
The scoring rubric
Score each question 0 to 2: 0 for a generic answer, 1 for a competent answer, 2 for a specific answer with numbers, named trade-offs, or an offer to look at your account. Out of 20:
- 16 to 20: an operator. Move to proposal and check the contract terms below.
- 10 to 15: competent but possibly package-driven. Probe the weakest answers again on a second call.
- Under 10: a sales motion. The account team behind it will not exceed the pitch.
Run the same rubric on at least two agencies so the scores mean something relative to each other, and weight questions 1, 3, and 10 double if you only remember three.
Red flags
- Open-rate reporting. Apple's Mail Privacy Protection inflates opens; agencies still selling open-rate wins are selling noise.
- Guaranteed revenue percentages. Attribution settings can manufacture any percentage without a dollar of real revenue.
- A proposal before an audit. Pricing without looking is packaging.
- No flow talk in the first call. Flows carry the durable revenue; opening with campaign volume optimizes the smaller lever.
- Twelve-month lock-ins with no performance checkpoints. Long contracts protect agencies from their own results; 60-to-90-day out clauses are standard.
- You do not own the account. Everything should be built in your Klaviyo under your login. Anything else is a hostage arrangement.
Green flags
- Published process. SOPs you can read beat portfolio screenshots, because process is what you are renting.
- An audit-first structure where the fix list stands alone whether or not you sign.
- Willingness to say who they are wrong for. An agency that names its bad-fit scenarios has a real specialization.
- Numbers with named windows. Every metric in the pitch deck should say which attribution window produced it.
- A paid test brief offer. One real email from your real assets beats every reference call.
Matching agency type to your situation
The rubric finds competent agencies; fit is a separate question. A production shop, a retention specialist, a full-funnel growth agency, and an enterprise media agency are different machines, and the most common failed engagement is a competent agency of the wrong shape. The 15 best Klaviyo agencies guide maps the shapes with scenario shortlists, and the in-house versus agency guide covers whether to hire an agency at all. For what an architecture-first engagement specifically looks like, see our retention agency page and services overview; the evaluation starting point either way is a proper account audit.
Frequently asked questions
How do I evaluate a Klaviyo agency before hiring?
Audit your own account first so you know the real problems, then ask ten specific questions covering attribution, flows, cadence, deliverability, vertical mechanics, and day-90 measurement, scoring answers on specificity. Operators answer with numbers and offers to look at your account; salespeople answer with case studies.
What questions should I ask a Klaviyo agency on the first call?
The three highest-signal ones: which attribution window do you report on and why, what would you fix in my account first (and can you look now), and how will we know at day 90 whether it worked. Full list of ten with model answers above.
What are the biggest red flags with email agencies?
Open-rate reporting, guaranteed revenue percentages, proposals before audits, no flow discussion in the first call, long lock-ins, and any arrangement where you do not own the Klaviyo account.
How long should I give an agency to show results?
Campaign improvements inside a month, architecture results by day 90 with attribution windows held constant. Judge flow rebuilds at the 90-day mark, not week two.
Should I run a paid test project before signing?
If two finalists both score well, yes: one paid brief from your real assets resolves a tie faster and more honestly than reference calls.