How does a managed pipeline like Shovels stack up to building your own database? We break down key considerations like cost, coverage, delivery, and more.
Building permit records are public, so in theory you can collect them yourself. For a single city or county, that may be the right approach. The challenge begins when you need coverage across dozens, hundreds, or even thousands of jurisdictions.
But getting the data is only part of the job. Every jurisdiction publishes permits differently. Records often arrive incomplete or duplicated, permit types are labeled inconsistently, and data pipelines can break whenever a city updates its website or portal.
That's where Shovels comes in. We handle the collection, normalization, and enrichment process across thousands of jurisdictions nationwide. We then deliver standardized permit data through an API, web application, or direct data warehouse share.
This page compares the two approaches so you can decide when it makes sense to build your own pipeline and when a managed dataset is the better use of time and resources.
KEY TAKEAWAYS
Collecting permit data is only the first step. Making it usable is where most of the effort goes.
Every jurisdiction publishes permits differently. Some have modern open-data portals, others rely on search forms, and many still distribute records through PDFs or public records requests.
Raw permit data also needs significant cleanup. Fields are inconsistent, duplicate records are common, and permit types vary widely from one jurisdiction to the next. A roofing permit might appear as "reroof," "roof - residential," or an internal code that means nothing outside that city. On top of that, data pipelines require ongoing maintenance as local governments update their websites and systems.
What starts as a simple data collection project can quickly become a long-term engineering commitment, whether it's handled by an internal team or outsourced to a contractor.
In short, Shovels is the intelligence layer for the built world. We collect permit data from jurisdictions across the country, and we use AI and quality-control workflows to turn messy municipal records into standardized, searchable data.
Each permit is enriched with categories, contractor information, parcel links, and other context that makes analysis easier. The result is a dataset that's ready to query, analyze, or load into your warehouse on day one.
Shovels also includes Decisions, a dataset that captures zoning, planning, and entitlement records. Together, permits and decisions data provide visibility into projects that are underway and projects still moving through the approval process.
Check out our Data Dictionary to learn more about the permit, contractor, decisions, and enhanced property data we offer.
Both approaches start with the same public records. The difference is in what happens next: collecting the data consistently, cleaning it, standardizing it, enriching it with additional context, and keeping it up to date over time.
The comparison below breaks down those tradeoffs across the areas that matter most, including coverage, data quality, enrichment, delivery options, and the ongoing cost of maintaining the data.
| Dimension | Doing it yourself (scraping) | Shovels |
|---|---|---|
| Coverage | ||
| Jurisdictions | One per scraper you build and maintain | 178M+ permits across 2,770+ jurisdictions |
| Working nationally | A new pipeline for every market | One nationwide dataset, strongest in major metros |
| Offline jurisdictions | Build direct relationships with each agency, request permit reports manually, and process offline formats yourself | Established jurisdiction relationships to source records that never go online |
| Data quality | ||
| Formatting | Each city's raw format; you normalize every field | Standardized, AI-cleaned fields and categories |
| Duplicates | You build and run dedup logic | Deduplicated automatically |
| Permit type tagging (solar, roofing, HVAC, EV) | Build and tune your own classifier | Pre-tagged with standardized categories |
| Enrichment | ||
| Contractor intelligence | Not included. Need to build it yourself | Permit history, inspection pass rates, build speed, revenue estimates, contacts |
| Parcel / property links | Manual joins you maintain | Linked to standardized parcel IDs and property data |
| Pre-permit signal | None | Decisions data tracks zoning and planning actions before permits file |
| Delivery | ||
| Access | Whatever you build | REST API, web app (Shovels Online), Snowflake / BigQuery / Databricks |
| Non-technical users | Needs a separate tool | Shovels Online with Charlie built in, no SQL required |
| Operating cost | ||
| Refresh | Whatever cadence you maintain | Core permit data refreshed twice monthly |
| Upkeep | Ongoing, pipeline may break when cities change their sites | Handled for you as sources change |
| Time to first usable data | Weeks to months | Same day |
Here's the bottom line: For a single city or a one-time pull, building it yourself is probably the more economical option. If you need many jurisdictions or historical data, a commercial provider is the better choice. Providers absorb the cost and labor of cleaning, enrichment, and ongoing maintenance, and they offer assistance when you need it.
An in-house pipeline covers only the jurisdictions a team integrates, and each new market adds another source to build and maintain. A commercial dataset typically provides national coverage already assembled. Shovels, for example, spans all 50 states across thousands of jurisdictions and most major metropolitan areas.
Not every agency publishes permits online. A meaningful share of records are only available through direct relationships with local jurisdictions, manual permit report requests, and offline formats like PDFs or public records requests. Building this yourself means establishing and maintaining those relationships agency by agency. Shovels already sources this data, including records that never appear on a city website.
Municipal records are rarely standardized, so turning them into a reliable dataset requires ongoing cleanup, deduplication, and normalization. If you build in-house, your team owns that process. With a provider like Shovels, the work is done for you and updated as jurisdictions change how they publish data.
Public permit records
Free, but raw
Do it yourself
Shovels
One managed step
Cleaned, deduped, enriched
Includes offline jurisdictions
Sourced via direct agency relationships
Delivered 3 ways
API
pull by geography & permit type
Shovels Online (Web App)
no SQL required
Shovels Enterprise
Snowflake · BigQuery · Databricks
Ready to query, day one
Weeks to months
and never finished
Same day
upkeep handled for you
Raw permit records provide limited context. Making them useful usually requires linking them to contractors, parcels, and related datasets, then enriching them with information such as contractor history, inspection outcomes, and parcel identifiers. That enrichment takes time and effort to build in-house. A commercial data provider usually offers enrichment out of the box.
Access methods determine how readily the data reaches a team's workflow. A commercial provider typically offers an API, a web interface, and warehouse delivery, whereas an in-house pipeline supports only the formats the team builds. Shovels, for example, offers three delivery methods: a web app, API, or a custom data feed for enterprise customers.
Because permits are public, the source data is free, but the full cost of an in-house approach includes engineering to build the pipeline, ongoing maintenance as sources change, and analyst time to clean and reconcile the output. A subscription consolidates those into a more predictable cost. Ultimately, balancing the cost depends on how many jurisdictions and how much enrichment a team requires.
How Shovels fits. Shovels applies the cleaning, standardization, and enrichment described above to permit, contractor, and property data from thousands of municipalities. Data delivery is available through the Shovels Online web application, the Shovels API, and Shovels Enterprise.
A DIY approach can absolutely make sense in certain situations. If you only need data from a single city or metro area, if you're doing a one-time analysis rather than building an ongoing data pipeline, or if you have engineering resources available and a well-defined set of requirements, collecting and managing the data yourself may be the more cost-effective option.
Shovels is designed for teams that need consistent coverage across many jurisdictions without having to build and maintain the underlying infrastructure. The trade-offs are worth understanding here as well.
Permit data is refreshed twice monthly, coverage is strongest in major metro areas (with some gaps in rural regions), and permits are not a perfect measure of completed construction since some projects are delayed, canceled, or never require a permit in the first place.
Ultimately, the decision comes down to whether you want to own the permit data pipeline or simply use the data.
If you're evaluating a DIY approach, look beyond the cost of collecting records. Account for the engineering time required to build and maintain scrapers, the analyst hours needed to clean and standardize the data, and the effort involved in adding enrichment and keeping everything current as jurisdictions change their systems.
Once those costs are fully loaded, compare them against a Shovels plan covering the same markets. For some teams, especially those focused on a small number of jurisdictions, building in-house is likely the right choice. For teams that need broad coverage, consistent data quality, and minimal maintenance overhead, a managed dataset is the better investment.
This is especially the case for enterprise teams, who often get more value from receiving permit data directly in their warehouse (Snowflake, BigQuery, or Databricks) and integrating it into existing workflows.
Finally, it is worth asking what your team is really for. If your edge comes from making decisions with good data, that is where your people should spend their time, not on manually collecting and maintaining it.
Have more questions about build vs buy? Explore the Shovels permit database for free, or talk to us directly about your needs.
Sign up free, pull the permits and contractors in a market you already know, and judge the coverage, deduplication, and enrichment for yourself.
A scraping team can pull permits, but you still own the cleaning, deduplication, standardization, and constant maintenance as cities change their sites, plus building any contractor or parcel enrichment yourself. Shovels delivers that as standardized, enriched data across 2,770+ jurisdictions, so you pay for usable intelligence rather than a pipeline you have to keep running. For offline jurisdictions, an offshore scraping team simply does not have the connections with local agencies needed to obtain records that never get published online. Shovels addresses this problem by building relationships with jurisdictions directly, ensuring we can access the valuable data you need.
A data dump is raw records in each city's own format. A permit intelligence platform standardizes those records into consistent fields and categories, deduplicates them, tags permit types, and links them to contractors and parcels, so the data is ready to analyze instead of ready to clean.
Yes, and for one city it can work. But even finding the right source is harder than it sounds. Is permitting handled by the city or the county? Does it depend on the permit type? Tracking down the correct portal for each jurisdiction is its own task before you write a line of scraping code. The difficulty scales with jurisdictions: every city has a different format and update method, and keeping many sources current and consistent is the part that turns a quick pull into an ongoing engineering project.
Shovels uses AI to clean and normalize inconsistent municipal text into standardized fields, deduplicate filings, and infer attributes like permit category. Field completeness still varies by jurisdiction, which is why Shovels publishes coverage detail rather than implying every record is complete.
Yes. Shovels offers a REST API to pull permit data by geography and permit type, a web app for non-technical users, and direct warehouse sharing into Snowflake, BigQuery, and Databricks.
You can join permits to parcels yourself, but it means sourcing and maintaining both datasets and reconciling addresses across them. Shovels links permits to standardized parcel IDs and property data out of the box, so the join is already done.