August 7, 2026

10 min read

Building a Custom AI Sales Prospecting Pipeline That Mapped 18 Million Square Feet

Why B2B sales teams outgrow off-the-shelf sales intelligence tools

Most B2B sales teams hit the same wall eventually: the ideal customer profile gets specific enough that a contact database can't filter for it anymore. ZoomInfo, Apollo, Clay, Demandbase, they're all built around the same core object, a company record enriched with firmographic and technographic fields (headcount, revenue band, tech stack, industry code) plus contact data layered on top. That works fine when your ICP is "Series B SaaS companies using Salesforce with 50-200 employees." It stops working the moment your ICP depends on something the vendor never collected: building square footage, hours of operation, utility rates by zip code, or visual evidence that a facility still has legacy lighting installed.

This isn't a tooling gap you can fix with a better filter or a smarter Clay recipe. It's a data-availability gap. No sales intelligence vendor scrapes county assessor records, building footprint data, or Street View imagery, because that data has nothing to do with the contact-and-company model their entire product is built on. If your growth motion depends on physical or operational characteristics of a facility, you're asking a phone-book to answer a satellite-imagery question.

What AI sales prospecting tools actually do well (and where they stop)

Give the category its due first. Modern AI sales prospecting tools are genuinely good at three things: identifying lead lists from firmographic and intent criteria, automating research and enrichment against existing contact and company databases, and personalizing outreach at scale. IBM frames this as AI analyzing sales data to reduce manual research and prioritize higher-probability accounts, and that's an accurate description of what Clay, Apollo, and Salesforce's own AI prospecting tools do (IBM, "AI Sales Prospecting"). Salesforce's positioning leans on the same three pillars: lead scoring, enrichment, and personalized sequencing (Salesforce).

Where they stop is data breadth, not intelligence. These platforms are excellent at recombining data they already have. Bombora's own explainer on B2B data enrichment lays out the standard taxonomy: firmographic (industry, size, revenue), technographic (installed software), and intent signals (content consumption, search behavior) (Bombora). Every one of those categories assumes the data already exists somewhere as a structured record a vendor can license or scrape. None of them touch physical-world attributes: how big a building is, how old its infrastructure looks, what it pays per kilowatt-hour. If your ICP needs that, you're not underusing your sales intelligence platform. You've reached the edge of what the category was built to do.

When your ideal customer profile needs data no vendor database has

The pattern shows up more often than people expect, and it's not limited to energy. Any company whose ICP depends on a physical or operational characteristic runs into this: facility square footage, equipment age, occupancy patterns, utility consumption, zoning, retrofit-readiness. Commercial energy efficiency firms need to know which buildings still run legacy lighting. Insurance and risk firms need building age and construction type. Industrial equipment resellers need facility size and current machinery vintage. In every case, the qualifying signal sits outside any contact database, scattered across county assessor sites, utility rate schedules, satellite and Street View imagery, and public filings.

The manual workaround is what most teams do first, and it's brutal. Someone visits a PE firm's portfolio page, Googles each portfolio company by hand, checks assessor records for square footage, and scrolls through Google Maps looking for visual clues about a facility's age. It takes weeks per pipeline, and by the time it's done, the data is already stale, deals close, portfolios change, rates shift quarterly. Scale that across even a few thousand target accounts and the math stops working: there aren't enough analyst-hours in the world.

Inside a custom account intelligence pipeline: the Lumen Global build

This is exactly the wall Lumen Global hit. Lumen partners with private equity firms, corporations, and REITs on commercial energy efficiency retrofits, and their pipeline depends on identifying which of a PE firm's portfolio companies operate facilities worth targeting. There are more than 30,000 PE firms in the US, each with a portfolio of companies, each with facilities that may or may not be retrofit candidates. Before Genta built their system, Lumen's team was doing the manual version above: site visits, manual searches, spreadsheets, no consistent scoring, and a process that took weeks and went stale almost immediately.

What we built instead needs only a PE firm name and website URL as input. From there, custom AI agents scrape the firm's site for portfolio company names, resolve facility addresses through the Google Maps Places API, and estimate building square footage using a multi-source approach that combines the Google Geocoding API, Overture Maps and Microsoft Building Footprint data, the Google Solar API, and a connected-component algorithm (built with Shapely) that avoids double-counting sprawling, multi-section industrial complexes as separate buildings. Electricity rates come from the EIA's public rate API, matched by location. And for the visual signal that used to require a human scrolling through Street View, the pipeline pulls Street View imagery and runs it through GPT-4o Vision, trained to flag wall-pack fixtures, parking lot pole types, and daytime lighting patterns, all proxies for legacy infrastructure that hasn't been retrofitted. Every facility comes out the other end with an automated 1-10 fit score.

The first run profiled more than 96 facilities and mapped over 18 million square feet of building footprint, a number that keeps growing as the pipeline runs. Adding a new PE firm to the system now takes seconds instead of weeks. Full details are in the Lumen Global case study, and the technical shape of it, agents orchestrating scraping, geocoding, computer vision, and scoring, is the same pattern we build for clients on our AI agents service line.

How much do B2B data enrichment tools cost, and when does a custom build pay for itself?

Off-the-shelf enrichment isn't cheap once you're past a small team. Autobound's pricing survey puts ZoomInfo's entry tier at roughly $14,900 or more per year, while Apollo starts around $49 per month for a much thinner feature set (Autobound, 2025 enrichment platform guide). Both are reasonable prices for what they do: appending firmographic and contact fields to records you already have. Neither includes a single field about building square footage or utility rates, because that was never the product.

The real comparison isn't tool cost versus build cost in isolation. It's tool cost versus the value of the deals you can't currently find. Lumen's typical deal size runs $500,000 to a few million dollars. Against that, one additional closed deal a year more than covers the cost of a custom pipeline, and the pipeline keeps surfacing new candidates indefinitely, at near-zero marginal cost per additional PE firm. That's a very different payback calculation than a per-seat SaaS subscription, and it's the same logic we walk through in more general terms in our piece on how to measure AI ROI and spot fake productivity: the question isn't "does this feel efficient," it's "what specific dollar outcome does this unlock, and how fast does it pay for itself."

A build-vs-buy framework for account intelligence

Before you commission a custom pipeline, run the decision through a few honest filters. This is roughly the checklist we use with clients during diagnosis:

  • Deal size and volume. If an average won deal is worth $50,000 or more, or your sales cycle is long enough that one missed target account represents real opportunity cost, custom enrichment usually pays for itself faster than people expect.

  • ICP specificity. If your qualifying criteria can be expressed as firmographic and technographic filters (industry, headcount, tech stack), buy. If your ICP depends on a physical, operational, or visual attribute no vendor tracks, buying gets you a partial answer at best.

  • Data source availability. Check whether the underlying data exists in a public or licensable form at all: assessor records, permit filings, utility rate schedules, satellite or Street View imagery, industry-specific registries. If it exists somewhere, even messily, it can usually be pipelined.

  • Refresh frequency. A one-time list-pull can be done manually or outsourced. A pipeline that needs to stay current as portfolios, rates, and facilities change is where automation earns its cost.

  • Team bandwidth. If your team is already spending analyst hours doing the manual version of this work, that cost is real, it's just hidden in payroll instead of a line-item invoice.

Most companies don't need to choose one path exclusively. A common pattern we see: keep a standard tool like Apollo or ZoomInfo for firmographic and contact enrichment on the broad funnel, and layer a custom pipeline on top for the narrow, high-value segment where the ICP gets specific. The tool handles volume; the custom system handles precision. This is the same underlying logic we cover in more general terms in our framework for buying versus building AI, and in the adjacent decision on the qualification side of the funnel in when to build a custom AI lead qualification agent instead of buying a SaaS seat.

What it actually takes to build this (timeline, team, and pitfalls)

A working prototype is not the same as a production pipeline, and this is where most build attempts underdeliver against the demo. The Lumen system touches at least six external data sources (PE firm sites, Google Places, Google Geocoding, Overture Maps, EIA rates, Street View plus vision analysis) and each one has its own rate limits, data gaps, and edge cases. Multi-building industrial complexes will get double-counted if your footprint logic doesn't dedupe connected structures. Street View coverage is inconsistent in rural and industrial zones. Vision models will confidently misclassify a fixture if you don't validate against a labeled sample first. None of this shows up in a two-week proof of concept; it shows up three months later when the pipeline has been running against real, messy data at volume.

Realistic timelines for a system like this run from a few weeks for a narrow first version to a few months for a fully productionized pipeline with monitoring, error handling, and a scoring model tuned against actual won deals. Budget for an iteration cycle where you validate the fit score against deals your sales team already knows are good or bad, because a scoring model that hasn't been checked against ground truth is just a plausible-looking number. And be honest about ownership: a vendor-built black box you can't modify defeats the purpose of going custom in the first place. If you build this, you should own the code, the data pipeline, and the scoring logic outright, not rent access to someone else's version of it. If you're evaluating who to build it with, our guide on what to ask before you hire an AI agent development company covers the specific questions that separate a real engineering partner from a demo shop.

If you're working through this decision, this is exactly what our Discovery phase maps out before any code gets written, and we're happy to compare notes.

Frequently asked questions

How is AI actually used in sales prospecting today?

Most AI sales prospecting today means three things: scoring leads against firmographic and intent criteria, automating research and data enrichment against existing databases, and personalizing outreach sequences at scale. Tools like Clay, Apollo, and Salesforce's AI prospecting features all work this way. The AI adds speed and scale to research that used to be manual, but it's still bounded by the data the underlying database already contains.

What's the difference between account-based marketing software and a custom-built account intelligence system?

ABM software (Demandbase, 6sense) targets accounts using firmographic, technographic, and intent data it already licenses or collects, evaluated against categories like the ones Gartner tracks for the ABM platform market (Gartner Reviews). A custom account intelligence system pulls from whatever data sources your specific ICP actually requires, including physical, operational, or visual signals no ABM vendor carries, and scores accounts against your own won-deal history rather than a generic model.

How much do B2B data enrichment tools like ZoomInfo, Apollo, or Clay cost, and when do they stop being worth it?

ZoomInfo starts around $14,900 per year and Apollo starts around $49 per month, per Autobound's 2025 pricing survey. They stop being worth it, on their own, once your ICP depends on data these platforms don't and can't carry, at which point you're paying for enrichment that doesn't actually enrich the fields that matter to your qualification.

Can AI identify sales targets based on physical or location-based criteria (building size, utility rates, equipment age)?

Yes, but not through standard sales intelligence platforms. It requires a custom pipeline combining geocoding APIs, building footprint data, public utility rate APIs, and computer vision analysis of satellite or Street View imagery. This is precisely the system Genta built for Lumen Global to identify retrofit-ready commercial facilities across PE portfolios.

How do you build a custom AI pipeline for account-based prospecting instead of buying a SaaS seat?

Start by mapping every data source your ICP criteria actually require, then chain APIs and agents that scrape, geocode, enrich, and score accounts against that specific data, validated against deals you've already won or lost. Expect a few weeks for a narrow first version and a few months for a monitored, production-grade system with a validated scoring model.

Tell us where the manual work hurts

We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.

Tell us where the manual work hurts

We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.

Tell us where the manual work hurts

We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.

August 7, 2026

10 min read

Building a Custom AI Sales Prospecting Pipeline That Mapped 18 Million Square Feet

Why B2B sales teams outgrow off-the-shelf sales intelligence tools

Most B2B sales teams hit the same wall eventually: the ideal customer profile gets specific enough that a contact database can't filter for it anymore. ZoomInfo, Apollo, Clay, Demandbase, they're all built around the same core object, a company record enriched with firmographic and technographic fields (headcount, revenue band, tech stack, industry code) plus contact data layered on top. That works fine when your ICP is "Series B SaaS companies using Salesforce with 50-200 employees." It stops working the moment your ICP depends on something the vendor never collected: building square footage, hours of operation, utility rates by zip code, or visual evidence that a facility still has legacy lighting installed.

This isn't a tooling gap you can fix with a better filter or a smarter Clay recipe. It's a data-availability gap. No sales intelligence vendor scrapes county assessor records, building footprint data, or Street View imagery, because that data has nothing to do with the contact-and-company model their entire product is built on. If your growth motion depends on physical or operational characteristics of a facility, you're asking a phone-book to answer a satellite-imagery question.

What AI sales prospecting tools actually do well (and where they stop)

Give the category its due first. Modern AI sales prospecting tools are genuinely good at three things: identifying lead lists from firmographic and intent criteria, automating research and enrichment against existing contact and company databases, and personalizing outreach at scale. IBM frames this as AI analyzing sales data to reduce manual research and prioritize higher-probability accounts, and that's an accurate description of what Clay, Apollo, and Salesforce's own AI prospecting tools do (IBM, "AI Sales Prospecting"). Salesforce's positioning leans on the same three pillars: lead scoring, enrichment, and personalized sequencing (Salesforce).

Where they stop is data breadth, not intelligence. These platforms are excellent at recombining data they already have. Bombora's own explainer on B2B data enrichment lays out the standard taxonomy: firmographic (industry, size, revenue), technographic (installed software), and intent signals (content consumption, search behavior) (Bombora). Every one of those categories assumes the data already exists somewhere as a structured record a vendor can license or scrape. None of them touch physical-world attributes: how big a building is, how old its infrastructure looks, what it pays per kilowatt-hour. If your ICP needs that, you're not underusing your sales intelligence platform. You've reached the edge of what the category was built to do.

When your ideal customer profile needs data no vendor database has

The pattern shows up more often than people expect, and it's not limited to energy. Any company whose ICP depends on a physical or operational characteristic runs into this: facility square footage, equipment age, occupancy patterns, utility consumption, zoning, retrofit-readiness. Commercial energy efficiency firms need to know which buildings still run legacy lighting. Insurance and risk firms need building age and construction type. Industrial equipment resellers need facility size and current machinery vintage. In every case, the qualifying signal sits outside any contact database, scattered across county assessor sites, utility rate schedules, satellite and Street View imagery, and public filings.

The manual workaround is what most teams do first, and it's brutal. Someone visits a PE firm's portfolio page, Googles each portfolio company by hand, checks assessor records for square footage, and scrolls through Google Maps looking for visual clues about a facility's age. It takes weeks per pipeline, and by the time it's done, the data is already stale, deals close, portfolios change, rates shift quarterly. Scale that across even a few thousand target accounts and the math stops working: there aren't enough analyst-hours in the world.

Inside a custom account intelligence pipeline: the Lumen Global build

This is exactly the wall Lumen Global hit. Lumen partners with private equity firms, corporations, and REITs on commercial energy efficiency retrofits, and their pipeline depends on identifying which of a PE firm's portfolio companies operate facilities worth targeting. There are more than 30,000 PE firms in the US, each with a portfolio of companies, each with facilities that may or may not be retrofit candidates. Before Genta built their system, Lumen's team was doing the manual version above: site visits, manual searches, spreadsheets, no consistent scoring, and a process that took weeks and went stale almost immediately.

What we built instead needs only a PE firm name and website URL as input. From there, custom AI agents scrape the firm's site for portfolio company names, resolve facility addresses through the Google Maps Places API, and estimate building square footage using a multi-source approach that combines the Google Geocoding API, Overture Maps and Microsoft Building Footprint data, the Google Solar API, and a connected-component algorithm (built with Shapely) that avoids double-counting sprawling, multi-section industrial complexes as separate buildings. Electricity rates come from the EIA's public rate API, matched by location. And for the visual signal that used to require a human scrolling through Street View, the pipeline pulls Street View imagery and runs it through GPT-4o Vision, trained to flag wall-pack fixtures, parking lot pole types, and daytime lighting patterns, all proxies for legacy infrastructure that hasn't been retrofitted. Every facility comes out the other end with an automated 1-10 fit score.

The first run profiled more than 96 facilities and mapped over 18 million square feet of building footprint, a number that keeps growing as the pipeline runs. Adding a new PE firm to the system now takes seconds instead of weeks. Full details are in the Lumen Global case study, and the technical shape of it, agents orchestrating scraping, geocoding, computer vision, and scoring, is the same pattern we build for clients on our AI agents service line.

How much do B2B data enrichment tools cost, and when does a custom build pay for itself?

Off-the-shelf enrichment isn't cheap once you're past a small team. Autobound's pricing survey puts ZoomInfo's entry tier at roughly $14,900 or more per year, while Apollo starts around $49 per month for a much thinner feature set (Autobound, 2025 enrichment platform guide). Both are reasonable prices for what they do: appending firmographic and contact fields to records you already have. Neither includes a single field about building square footage or utility rates, because that was never the product.

The real comparison isn't tool cost versus build cost in isolation. It's tool cost versus the value of the deals you can't currently find. Lumen's typical deal size runs $500,000 to a few million dollars. Against that, one additional closed deal a year more than covers the cost of a custom pipeline, and the pipeline keeps surfacing new candidates indefinitely, at near-zero marginal cost per additional PE firm. That's a very different payback calculation than a per-seat SaaS subscription, and it's the same logic we walk through in more general terms in our piece on how to measure AI ROI and spot fake productivity: the question isn't "does this feel efficient," it's "what specific dollar outcome does this unlock, and how fast does it pay for itself."

A build-vs-buy framework for account intelligence

Before you commission a custom pipeline, run the decision through a few honest filters. This is roughly the checklist we use with clients during diagnosis:

  • Deal size and volume. If an average won deal is worth $50,000 or more, or your sales cycle is long enough that one missed target account represents real opportunity cost, custom enrichment usually pays for itself faster than people expect.

  • ICP specificity. If your qualifying criteria can be expressed as firmographic and technographic filters (industry, headcount, tech stack), buy. If your ICP depends on a physical, operational, or visual attribute no vendor tracks, buying gets you a partial answer at best.

  • Data source availability. Check whether the underlying data exists in a public or licensable form at all: assessor records, permit filings, utility rate schedules, satellite or Street View imagery, industry-specific registries. If it exists somewhere, even messily, it can usually be pipelined.

  • Refresh frequency. A one-time list-pull can be done manually or outsourced. A pipeline that needs to stay current as portfolios, rates, and facilities change is where automation earns its cost.

  • Team bandwidth. If your team is already spending analyst hours doing the manual version of this work, that cost is real, it's just hidden in payroll instead of a line-item invoice.

Most companies don't need to choose one path exclusively. A common pattern we see: keep a standard tool like Apollo or ZoomInfo for firmographic and contact enrichment on the broad funnel, and layer a custom pipeline on top for the narrow, high-value segment where the ICP gets specific. The tool handles volume; the custom system handles precision. This is the same underlying logic we cover in more general terms in our framework for buying versus building AI, and in the adjacent decision on the qualification side of the funnel in when to build a custom AI lead qualification agent instead of buying a SaaS seat.

What it actually takes to build this (timeline, team, and pitfalls)

A working prototype is not the same as a production pipeline, and this is where most build attempts underdeliver against the demo. The Lumen system touches at least six external data sources (PE firm sites, Google Places, Google Geocoding, Overture Maps, EIA rates, Street View plus vision analysis) and each one has its own rate limits, data gaps, and edge cases. Multi-building industrial complexes will get double-counted if your footprint logic doesn't dedupe connected structures. Street View coverage is inconsistent in rural and industrial zones. Vision models will confidently misclassify a fixture if you don't validate against a labeled sample first. None of this shows up in a two-week proof of concept; it shows up three months later when the pipeline has been running against real, messy data at volume.

Realistic timelines for a system like this run from a few weeks for a narrow first version to a few months for a fully productionized pipeline with monitoring, error handling, and a scoring model tuned against actual won deals. Budget for an iteration cycle where you validate the fit score against deals your sales team already knows are good or bad, because a scoring model that hasn't been checked against ground truth is just a plausible-looking number. And be honest about ownership: a vendor-built black box you can't modify defeats the purpose of going custom in the first place. If you build this, you should own the code, the data pipeline, and the scoring logic outright, not rent access to someone else's version of it. If you're evaluating who to build it with, our guide on what to ask before you hire an AI agent development company covers the specific questions that separate a real engineering partner from a demo shop.

If you're working through this decision, this is exactly what our Discovery phase maps out before any code gets written, and we're happy to compare notes.

Frequently asked questions

How is AI actually used in sales prospecting today?

Most AI sales prospecting today means three things: scoring leads against firmographic and intent criteria, automating research and data enrichment against existing databases, and personalizing outreach sequences at scale. Tools like Clay, Apollo, and Salesforce's AI prospecting features all work this way. The AI adds speed and scale to research that used to be manual, but it's still bounded by the data the underlying database already contains.

What's the difference between account-based marketing software and a custom-built account intelligence system?

ABM software (Demandbase, 6sense) targets accounts using firmographic, technographic, and intent data it already licenses or collects, evaluated against categories like the ones Gartner tracks for the ABM platform market (Gartner Reviews). A custom account intelligence system pulls from whatever data sources your specific ICP actually requires, including physical, operational, or visual signals no ABM vendor carries, and scores accounts against your own won-deal history rather than a generic model.

How much do B2B data enrichment tools like ZoomInfo, Apollo, or Clay cost, and when do they stop being worth it?

ZoomInfo starts around $14,900 per year and Apollo starts around $49 per month, per Autobound's 2025 pricing survey. They stop being worth it, on their own, once your ICP depends on data these platforms don't and can't carry, at which point you're paying for enrichment that doesn't actually enrich the fields that matter to your qualification.

Can AI identify sales targets based on physical or location-based criteria (building size, utility rates, equipment age)?

Yes, but not through standard sales intelligence platforms. It requires a custom pipeline combining geocoding APIs, building footprint data, public utility rate APIs, and computer vision analysis of satellite or Street View imagery. This is precisely the system Genta built for Lumen Global to identify retrofit-ready commercial facilities across PE portfolios.

How do you build a custom AI pipeline for account-based prospecting instead of buying a SaaS seat?

Start by mapping every data source your ICP criteria actually require, then chain APIs and agents that scrape, geocode, enrich, and score accounts against that specific data, validated against deals you've already won or lost. Expect a few weeks for a narrow first version and a few months for a monitored, production-grade system with a validated scoring model.

Tell us where the manual work hurts

We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.

Tell us where the manual work hurts

We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.

Tell us where the manual work hurts

We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.

August 7, 2026

10 min read

Building a Custom AI Sales Prospecting Pipeline That Mapped 18 Million Square Feet

Why B2B sales teams outgrow off-the-shelf sales intelligence tools

Most B2B sales teams hit the same wall eventually: the ideal customer profile gets specific enough that a contact database can't filter for it anymore. ZoomInfo, Apollo, Clay, Demandbase, they're all built around the same core object, a company record enriched with firmographic and technographic fields (headcount, revenue band, tech stack, industry code) plus contact data layered on top. That works fine when your ICP is "Series B SaaS companies using Salesforce with 50-200 employees." It stops working the moment your ICP depends on something the vendor never collected: building square footage, hours of operation, utility rates by zip code, or visual evidence that a facility still has legacy lighting installed.

This isn't a tooling gap you can fix with a better filter or a smarter Clay recipe. It's a data-availability gap. No sales intelligence vendor scrapes county assessor records, building footprint data, or Street View imagery, because that data has nothing to do with the contact-and-company model their entire product is built on. If your growth motion depends on physical or operational characteristics of a facility, you're asking a phone-book to answer a satellite-imagery question.

What AI sales prospecting tools actually do well (and where they stop)

Give the category its due first. Modern AI sales prospecting tools are genuinely good at three things: identifying lead lists from firmographic and intent criteria, automating research and enrichment against existing contact and company databases, and personalizing outreach at scale. IBM frames this as AI analyzing sales data to reduce manual research and prioritize higher-probability accounts, and that's an accurate description of what Clay, Apollo, and Salesforce's own AI prospecting tools do (IBM, "AI Sales Prospecting"). Salesforce's positioning leans on the same three pillars: lead scoring, enrichment, and personalized sequencing (Salesforce).

Where they stop is data breadth, not intelligence. These platforms are excellent at recombining data they already have. Bombora's own explainer on B2B data enrichment lays out the standard taxonomy: firmographic (industry, size, revenue), technographic (installed software), and intent signals (content consumption, search behavior) (Bombora). Every one of those categories assumes the data already exists somewhere as a structured record a vendor can license or scrape. None of them touch physical-world attributes: how big a building is, how old its infrastructure looks, what it pays per kilowatt-hour. If your ICP needs that, you're not underusing your sales intelligence platform. You've reached the edge of what the category was built to do.

When your ideal customer profile needs data no vendor database has

The pattern shows up more often than people expect, and it's not limited to energy. Any company whose ICP depends on a physical or operational characteristic runs into this: facility square footage, equipment age, occupancy patterns, utility consumption, zoning, retrofit-readiness. Commercial energy efficiency firms need to know which buildings still run legacy lighting. Insurance and risk firms need building age and construction type. Industrial equipment resellers need facility size and current machinery vintage. In every case, the qualifying signal sits outside any contact database, scattered across county assessor sites, utility rate schedules, satellite and Street View imagery, and public filings.

The manual workaround is what most teams do first, and it's brutal. Someone visits a PE firm's portfolio page, Googles each portfolio company by hand, checks assessor records for square footage, and scrolls through Google Maps looking for visual clues about a facility's age. It takes weeks per pipeline, and by the time it's done, the data is already stale, deals close, portfolios change, rates shift quarterly. Scale that across even a few thousand target accounts and the math stops working: there aren't enough analyst-hours in the world.

Inside a custom account intelligence pipeline: the Lumen Global build

This is exactly the wall Lumen Global hit. Lumen partners with private equity firms, corporations, and REITs on commercial energy efficiency retrofits, and their pipeline depends on identifying which of a PE firm's portfolio companies operate facilities worth targeting. There are more than 30,000 PE firms in the US, each with a portfolio of companies, each with facilities that may or may not be retrofit candidates. Before Genta built their system, Lumen's team was doing the manual version above: site visits, manual searches, spreadsheets, no consistent scoring, and a process that took weeks and went stale almost immediately.

What we built instead needs only a PE firm name and website URL as input. From there, custom AI agents scrape the firm's site for portfolio company names, resolve facility addresses through the Google Maps Places API, and estimate building square footage using a multi-source approach that combines the Google Geocoding API, Overture Maps and Microsoft Building Footprint data, the Google Solar API, and a connected-component algorithm (built with Shapely) that avoids double-counting sprawling, multi-section industrial complexes as separate buildings. Electricity rates come from the EIA's public rate API, matched by location. And for the visual signal that used to require a human scrolling through Street View, the pipeline pulls Street View imagery and runs it through GPT-4o Vision, trained to flag wall-pack fixtures, parking lot pole types, and daytime lighting patterns, all proxies for legacy infrastructure that hasn't been retrofitted. Every facility comes out the other end with an automated 1-10 fit score.

The first run profiled more than 96 facilities and mapped over 18 million square feet of building footprint, a number that keeps growing as the pipeline runs. Adding a new PE firm to the system now takes seconds instead of weeks. Full details are in the Lumen Global case study, and the technical shape of it, agents orchestrating scraping, geocoding, computer vision, and scoring, is the same pattern we build for clients on our AI agents service line.

How much do B2B data enrichment tools cost, and when does a custom build pay for itself?

Off-the-shelf enrichment isn't cheap once you're past a small team. Autobound's pricing survey puts ZoomInfo's entry tier at roughly $14,900 or more per year, while Apollo starts around $49 per month for a much thinner feature set (Autobound, 2025 enrichment platform guide). Both are reasonable prices for what they do: appending firmographic and contact fields to records you already have. Neither includes a single field about building square footage or utility rates, because that was never the product.

The real comparison isn't tool cost versus build cost in isolation. It's tool cost versus the value of the deals you can't currently find. Lumen's typical deal size runs $500,000 to a few million dollars. Against that, one additional closed deal a year more than covers the cost of a custom pipeline, and the pipeline keeps surfacing new candidates indefinitely, at near-zero marginal cost per additional PE firm. That's a very different payback calculation than a per-seat SaaS subscription, and it's the same logic we walk through in more general terms in our piece on how to measure AI ROI and spot fake productivity: the question isn't "does this feel efficient," it's "what specific dollar outcome does this unlock, and how fast does it pay for itself."

A build-vs-buy framework for account intelligence

Before you commission a custom pipeline, run the decision through a few honest filters. This is roughly the checklist we use with clients during diagnosis:

  • Deal size and volume. If an average won deal is worth $50,000 or more, or your sales cycle is long enough that one missed target account represents real opportunity cost, custom enrichment usually pays for itself faster than people expect.

  • ICP specificity. If your qualifying criteria can be expressed as firmographic and technographic filters (industry, headcount, tech stack), buy. If your ICP depends on a physical, operational, or visual attribute no vendor tracks, buying gets you a partial answer at best.

  • Data source availability. Check whether the underlying data exists in a public or licensable form at all: assessor records, permit filings, utility rate schedules, satellite or Street View imagery, industry-specific registries. If it exists somewhere, even messily, it can usually be pipelined.

  • Refresh frequency. A one-time list-pull can be done manually or outsourced. A pipeline that needs to stay current as portfolios, rates, and facilities change is where automation earns its cost.

  • Team bandwidth. If your team is already spending analyst hours doing the manual version of this work, that cost is real, it's just hidden in payroll instead of a line-item invoice.

Most companies don't need to choose one path exclusively. A common pattern we see: keep a standard tool like Apollo or ZoomInfo for firmographic and contact enrichment on the broad funnel, and layer a custom pipeline on top for the narrow, high-value segment where the ICP gets specific. The tool handles volume; the custom system handles precision. This is the same underlying logic we cover in more general terms in our framework for buying versus building AI, and in the adjacent decision on the qualification side of the funnel in when to build a custom AI lead qualification agent instead of buying a SaaS seat.

What it actually takes to build this (timeline, team, and pitfalls)

A working prototype is not the same as a production pipeline, and this is where most build attempts underdeliver against the demo. The Lumen system touches at least six external data sources (PE firm sites, Google Places, Google Geocoding, Overture Maps, EIA rates, Street View plus vision analysis) and each one has its own rate limits, data gaps, and edge cases. Multi-building industrial complexes will get double-counted if your footprint logic doesn't dedupe connected structures. Street View coverage is inconsistent in rural and industrial zones. Vision models will confidently misclassify a fixture if you don't validate against a labeled sample first. None of this shows up in a two-week proof of concept; it shows up three months later when the pipeline has been running against real, messy data at volume.

Realistic timelines for a system like this run from a few weeks for a narrow first version to a few months for a fully productionized pipeline with monitoring, error handling, and a scoring model tuned against actual won deals. Budget for an iteration cycle where you validate the fit score against deals your sales team already knows are good or bad, because a scoring model that hasn't been checked against ground truth is just a plausible-looking number. And be honest about ownership: a vendor-built black box you can't modify defeats the purpose of going custom in the first place. If you build this, you should own the code, the data pipeline, and the scoring logic outright, not rent access to someone else's version of it. If you're evaluating who to build it with, our guide on what to ask before you hire an AI agent development company covers the specific questions that separate a real engineering partner from a demo shop.

If you're working through this decision, this is exactly what our Discovery phase maps out before any code gets written, and we're happy to compare notes.

Frequently asked questions

How is AI actually used in sales prospecting today?

Most AI sales prospecting today means three things: scoring leads against firmographic and intent criteria, automating research and data enrichment against existing databases, and personalizing outreach sequences at scale. Tools like Clay, Apollo, and Salesforce's AI prospecting features all work this way. The AI adds speed and scale to research that used to be manual, but it's still bounded by the data the underlying database already contains.

What's the difference between account-based marketing software and a custom-built account intelligence system?

ABM software (Demandbase, 6sense) targets accounts using firmographic, technographic, and intent data it already licenses or collects, evaluated against categories like the ones Gartner tracks for the ABM platform market (Gartner Reviews). A custom account intelligence system pulls from whatever data sources your specific ICP actually requires, including physical, operational, or visual signals no ABM vendor carries, and scores accounts against your own won-deal history rather than a generic model.

How much do B2B data enrichment tools like ZoomInfo, Apollo, or Clay cost, and when do they stop being worth it?

ZoomInfo starts around $14,900 per year and Apollo starts around $49 per month, per Autobound's 2025 pricing survey. They stop being worth it, on their own, once your ICP depends on data these platforms don't and can't carry, at which point you're paying for enrichment that doesn't actually enrich the fields that matter to your qualification.

Can AI identify sales targets based on physical or location-based criteria (building size, utility rates, equipment age)?

Yes, but not through standard sales intelligence platforms. It requires a custom pipeline combining geocoding APIs, building footprint data, public utility rate APIs, and computer vision analysis of satellite or Street View imagery. This is precisely the system Genta built for Lumen Global to identify retrofit-ready commercial facilities across PE portfolios.

How do you build a custom AI pipeline for account-based prospecting instead of buying a SaaS seat?

Start by mapping every data source your ICP criteria actually require, then chain APIs and agents that scrape, geocode, enrich, and score accounts against that specific data, validated against deals you've already won or lost. Expect a few weeks for a narrow first version and a few months for a monitored, production-grade system with a validated scoring model.

Tell us where the manual work hurts

We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.

Tell us where the manual work hurts

We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.

Tell us where the manual work hurts

We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.