MICROMARKETING Book With Tony
← AI & Automation

Use Breeze Data Studio and Cortex Analyst Before Q4 Budget Planning

September is when campaign influence gets expensive. A $68,000 opportunity can look paid-search sourced in HubSpot, webinar-influenced in a campaign report, and organic-assisted in a content dashboard. Then Q4 planning starts, and three channel owners walk into the same meeting with the same revenue. I would not try to settle that fight inside a spreadsheet. […]

September is when campaign influence gets expensive. A $68,000 opportunity can look paid-search sourced in HubSpot, webinar-influenced in a campaign report, and organic-assisted in a content dashboard. Then Q4 planning starts, and three channel owners walk into the same meeting with the same revenue.

I would not try to settle that fight inside a spreadsheet. HubSpot’s Fall 2025 launch put Data Studio inside Data Hub, with AI-assisted tools for combining and enhancing customer data from sources like Google Sheets, Excel files, and Snowflake. Snowflake’s Cortex Analyst handles the other half of the problem: governed natural-language questions over structured data in Snowflake, using semantic views or semantic models so “booked ARR” means the same thing every time.

That pairing matters before Q4 budget planning because campaign influence is rarely wrong in one dramatic way. It is wrong in twelve dull ways. Duplicate campaign names. Stale lifecycle stages. UTM values with three casing styles. HubSpot deal amounts that never matched booked revenue. Contacts attached to closed-won deals while still sitting in Lead.

Clean the CRM records in HubSpot. Then make Snowflake prove the money.

Why This Breaks Before Q4

A founder running paid and organic with a lean marketing operator might spend $42,000 in Q3 across Google Ads, LinkedIn, founder posts, two newsletters, and a late-August webinar. HubSpot has contacts, companies, deals, lifecycle stages, campaigns, UTMs, ad IDs, and source fields. Snowflake has invoices, subscriptions, finance-approved ARR, refunds, product usage, segment, and renewal status.

Then the Q4 question lands: where should the next $90,000 go?

The first answer usually looks too clean. LinkedIn influenced $310,000 in pipeline. Organic influenced $270,000. The webinar influenced $180,000. Google non-brand influenced $220,000. Total Q3 created pipeline was only $640,000, so the math is already telling you the attribution model is handing out duplicate credit.

This is the point where operators get blamed for data nobody designed for budget review. The campaign report was built for activity tracking. The finance table was built for revenue reporting. The board slide wants a decision by channel, and neither system can answer that alone.

What Breeze Data Studio Should Fix First

HubSpot says Data Studio can combine data in a spreadsheet-like interface, connect external sources, and create curated datasets that flow into CRM lists, workflows, and reports. That is useful because the first pass should happen close to the CRM mess.

Start with campaign identity. In one portal I would expect to see variants like q3-ai-webinar, Q3_AI_Webinar, 2026-09-ai-readiness, and ai-readiness-webinar-sept. Pick one canonical value, such as 2026-q3-ai-readiness-webinar, and map the variants to it. Keep the raw UTM value. You will need it when somebody asks why a record moved.

Then normalize channel. I use a short list for teams under $20 million ARR: paid search, paid social, organic search, organic social, partner, event, lifecycle, outbound, direct, unknown. Do not turn unknown into a fake answer. A blank field is annoying. A guessed source can survive into three dashboards and burn a budget call in November.

Next, repair lifecycle drift. If a contact is tied to a closed-won deal dated August 18, 2026, but the lifecycle stage still says Lead, that is not a marketing insight. It is a workflow failure, a sync issue, or a property mapping problem. Breeze Data Studio and HubSpot’s 2025 Data Quality tools are built for exactly this kind of cleanup, including duplicates, missing information, inconsistent formats, and bad customer data.

I would also create a cleanup audit field. Call it data_studio_rule_id or crm_cleanup_batch. Use values like 2026-09-q4-planning-pass-01. Six weeks later, when a sales leader asks why 41 webinar touches became 28 influenced opportunities, that field keeps the conversation grounded.

What Cortex Analyst Should Challenge

Snowflake describes Cortex Analyst as a managed Cortex feature that lets business users ask natural-language questions over structured data and receive answers without writing SQL. The important part for budget work is the semantic layer. Snowflake now recommends Semantic Views for new implementations, while legacy YAML semantic models are still supported.

That means you can define the business terms before people start asking questions. Booked ARR should come from finance-owned subscription or invoice tables. Influenced pipeline should use weighted deal influence, not raw campaign membership. Opportunity created date should come from the CRM opportunity object. Late-stage assist should mean the campaign touch happened after opportunity creation but before close.

Cortex Analyst is not there to make the budget decision. It is there to make the cleaned HubSpot data answer finance-grade questions.

I would load a deal-level influence table into Snowflake with these fields: deal_id, company_id, contact_id, canonical_campaign, canonical_channel, touch_date, influence_role, influence_weight, hubspot_amount, cleaned_lifecycle_stage, raw_utm_campaign, and data_studio_rule_id.

Then join it to finance fields: booked_arr, recognized_revenue, invoice_status, refund_amount, subscription_start_date, segment, and sales_motion. If HubSpot says a deal is worth $72,000 and finance booked $54,000, do not blend the numbers. Flag the mismatch and use finance for planning.

The Reconciliation Runbook

Run the first pass by September 15 for a normal Q4 planning cycle. If quarter-end is September 30, that gives you two weeks to fix field mappings, inspect weird records, and rerun the questions before the deck freezes.

In HubSpot, freeze the Q3 campaign universe. Export campaign memberships, contacts, companies, deals, lifecycle stages, original source, latest source, UTM fields, HubSpot campaign IDs, Google Ads campaign IDs, and LinkedIn campaign group IDs. Names lie. IDs lie less.

Use Breeze Data Studio to create the cleaned campaign dataset. Map UTM variants. Standardize channel. Split first-touch, conversion-touch, opportunity-touch, and late-stage assist. For a simple B2B model, I would start with 40 percent credit for first touch, 30 percent for lead conversion, 20 percent for opportunity creation, and 10 percent for late-stage assist. It will not satisfy every attribution purist. It will stop one deal from being counted four times.

Load the cleaned dataset into Snowflake. Build or update the Semantic View so Cortex Analyst knows which tables carry campaign influence, which table owns booked ARR, and which date fields matter. Include synonyms operators actually use, such as “sourced pipeline,” “influenced ARR,” “late assist,” and “Q3 webinars.”

Then ask blunt questions:

  • Which opportunities have more than one campaign claiming full influence?
  • Which closed-won deals have contacts still marked Lead or MQL?
  • Which campaigns show a HubSpot amount more than 10 percent above booked ARR?
  • Which campaign touches happened after opportunity creation but were counted as sourced?
  • Which channels created pipeline in Q3 but produced no closed-won ARR within 90 days?

Those questions catch the records that move money. A duplicated $120,000 enterprise deal can distort paid social. A stale lifecycle workflow can make paid acquisition look broken. A 28 percent HubSpot-to-finance revenue gap can make a partner campaign look better than it was.

What Goes Into The Budget Packet

The budget packet should not include a 19-tab attribution workbook. Put the reconciliation on one page and keep the definitions tight.

Use five headline numbers: deduplicated influenced pipeline, reconciled booked ARR, duplicate-influence opportunity count, stale lifecycle-stage contact count, and revenue mismatch dollars above the 10 percent threshold.

Then show channel performance after reconciliation. For example, paid search influenced $312,000 in deduped pipeline and $86,000 in booked ARR. Organic search influenced $184,000 in pipeline and $91,000 in booked ARR. The September webinar influenced $140,000 in pipeline, but $58,000 came from opportunities already open before registration, so those touches were treated as late-stage assist.

That last sentence is the work. Nobody needs the model to sound clever. The operator needs to explain why the webinar budget goes from $18,000 to $10,000 while organic refresh work goes from $7,500 to $12,000.

I would include three notes in plain English:

“Google Ads non-brand produced 19 opportunities and $86,000 booked ARR. Four deals had HubSpot amounts at least 20 percent above finance-booked ARR, so finance values were used.”

“The AI readiness webinar touched 14 opportunities. Six were already open before registration, so they were counted as late-stage assist at 10 percent influence weight.”

“Organic comparison pages influenced $91,000 booked ARR with no duplicate campaign claims above the 10 percent threshold. Increase Q4 content refresh budget by $4,500.”

That is enough for a founder to make a call. It is also enough for a marketing operator to defend the call when sales asks why a favorite campaign lost credit.

Where Teams Get This Wrong

The first mistake is treating HubSpot cleanup as revenue truth. Breeze Data Studio can make CRM records usable, but finance still owns booked ARR, refunds, subscription start dates, and revenue recognition.

The second mistake is asking Cortex Analyst budget questions before the semantic layer is ready. Natural language over vague metrics gives you fast confusion. Define the metrics first, especially influenced_pipeline, booked_arr, first_touch_date, opportunity_created_date, and late_stage_assist.

The third mistake is overwriting raw data. Keep raw campaign names, raw UTMs, original lifecycle stages, and the cleanup rule that changed them. Attribution arguments become shorter when the audit trail is visible.

The fourth mistake is giving every touch equal credit because it feels politically easier. Equal credit rewards noisy journeys. A founder LinkedIn post, a Google ad click, and a partner webinar can all matter, but they should not each receive 100 percent of the same deal.

The operating rule is simple: use HubSpot Breeze Data Studio to clean and map the campaign records, then use Snowflake Cortex Analyst to test those records against governed revenue data. Run that before the Q4 budget meeting, not after the spend shifts.

AI automation earns its keep when it catches expensive mistakes early. Duplicate influence, stale lifecycle stages, and revenue mismatches are not glamorous problems. They are exactly the problems that move next quarter’s money.