Shovels Shovels - Building permit and contractor intelligence
COMPARE

Building permit data: Shovels vs. scraping it yourself

How does a managed pipeline like Shovels stack up to building your own database? We break down key considerations like cost, coverage, delivery, and more.

Building permit records are public, so in theory you can collect them yourself. For a single city or county, that may be the right approach. The challenge begins when you need coverage across dozens, hundreds, or even thousands of jurisdictions.

But getting the data is only part of the job. Every jurisdiction publishes permits differently. Records often arrive incomplete or duplicated, permit types are labeled inconsistently, and data pipelines can break whenever a city updates its website or portal.

That's where Shovels comes in. We handle the collection, normalization, and enrichment process across thousands of jurisdictions nationwide. We then deliver standardized permit data through an API, web application, or direct data warehouse share.

This page compares the two approaches so you can decide when it makes sense to build your own pipeline and when a managed dataset is the better use of time and resources.

KEY TAKEAWAYS

  • The cost of DIY is cleaning, not access. Pulling permits is easy but standardizing, deduplicating, and maintaining them across cities is the real work.
  • Scaling coverage can be costly and inefficient. One scraper only covers one city. Shovels covers 178M+ permits across 2,770+ jurisdictions nationwide.
  • Raw permits lack context. Shovels links each permit to contractors, city decisions, and parcels, while tagging permit types automatically.
  • "Free" public data is a recurring cost. Engineering time, maintenance, and analyst cleanup add up fast.
  • Shovels turns public records into Shovel-ready intelligence, delivered via API, web app (Shovels Online), or data warehouse.

What does scraping building permit data yourself actually involve?

Collecting permit data is only the first step. Making it usable is where most of the effort goes.

Every jurisdiction publishes permits differently. Some have modern open-data portals, others rely on search forms, and many still distribute records through PDFs or public records requests.

Raw permit data also needs significant cleanup. Fields are inconsistent, duplicate records are common, and permit types vary widely from one jurisdiction to the next. A roofing permit might appear as "reroof," "roof - residential," or an internal code that means nothing outside that city. On top of that, data pipelines require ongoing maintenance as local governments update their websites and systems.

What starts as a simple data collection project can quickly become a long-term engineering commitment, whether it's handled by an internal team or outsourced to a contractor.

What is Shovels?

In short, Shovels is the intelligence layer for the built world. We collect permit data from jurisdictions across the country, and we use AI and quality-control workflows to turn messy municipal records into standardized, searchable data.

Each permit is enriched with categories, contractor information, parcel links, and other context that makes analysis easier. The result is a dataset that's ready to query, analyze, or load into your warehouse on day one.

Shovels also includes Decisions, a dataset that captures zoning, planning, and entitlement records. Together, permits and decisions data provide visibility into projects that are underway and projects still moving through the approval process.

Check out our Data Dictionary to learn more about the permit, contractor, decisions, and enhanced property data we offer.

How does DIY permit scraping compare to Shovels?

Both approaches start with the same public records. The difference is in what happens next: collecting the data consistently, cleaning it, standardizing it, enriching it with additional context, and keeping it up to date over time.

The comparison below breaks down those tradeoffs across the areas that matter most, including coverage, data quality, enrichment, delivery options, and the ongoing cost of maintaining the data.

Dimension Doing it yourself (scraping) Shovels
Coverage
Jurisdictions One per scraper you build and maintain 178M+ permits across 2,770+ jurisdictions
Working nationally A new pipeline for every market One nationwide dataset, strongest in major metros
Offline jurisdictions Build direct relationships with each agency, request permit reports manually, and process offline formats yourself Established jurisdiction relationships to source records that never go online
Data quality
Formatting Each city's raw format; you normalize every field Standardized, AI-cleaned fields and categories
Duplicates You build and run dedup logic Deduplicated automatically
Permit type tagging (solar, roofing, HVAC, EV) Build and tune your own classifier Pre-tagged with standardized categories
Enrichment
Contractor intelligence Not included. Need to build it yourself Permit history, inspection pass rates, build speed, revenue estimates, contacts
Parcel / property links Manual joins you maintain Linked to standardized parcel IDs and property data
Pre-permit signal None Decisions data tracks zoning and planning actions before permits file
Delivery
Access Whatever you build REST API, web app (Shovels Online), Snowflake / BigQuery / Databricks
Non-technical users Needs a separate tool Shovels Online with Charlie built in, no SQL required
Operating cost
Refresh Whatever cadence you maintain Core permit data refreshed twice monthly
Upkeep Ongoing, pipeline may break when cities change their sites Handled for you as sources change
Time to first usable data Weeks to months Same day

Here's the bottom line: For a single city or a one-time pull, building it yourself is probably the more economical option. If you need many jurisdictions or historical data, a commercial provider is the better choice. Providers absorb the cost and labor of cleaning, enrichment, and ongoing maintenance, and they offer assistance when you need it.

The differences that matter most

  1. 01

    Coverage

    An in-house pipeline covers only the jurisdictions a team integrates, and each new market adds another source to build and maintain. A commercial dataset typically provides national coverage already assembled. Shovels, for example, spans all 50 states across thousands of jurisdictions and most major metropolitan areas.

  2. 02

    Offline jurisdictions

    Not every agency publishes permits online. A meaningful share of records are only available through direct relationships with local jurisdictions, manual permit report requests, and offline formats like PDFs or public records requests. Building this yourself means establishing and maintaining those relationships agency by agency. Shovels already sources this data, including records that never appear on a city website.

  3. 03

    Cleaning and standardization

    Municipal records are rarely standardized, so turning them into a reliable dataset requires ongoing cleanup, deduplication, and normalization. If you build in-house, your team owns that process. With a provider like Shovels, the work is done for you and updated as jurisdictions change how they publish data.

    Public permit records

    Free, but raw

    Do it yourself

    • Find the right portal
    • Scrape the records
    • Clean & deduplicate
    • Standardize & tag
    • Enrich with contractors & parcels
    • Build & maintain delivery
    • Offline records: no portal exists

    Shovels

    One managed step

    Cleaned, deduped, enriched
    Includes offline jurisdictions
    Sourced via direct agency relationships

    Delivered 3 ways

    • API

      pull by geography & permit type

    • Shovels Online (Web App)

      no SQL required

    • Shovels Enterprise

      Snowflake · BigQuery · Databricks

    Ready to query, day one

    Weeks to months

    and never finished

    Same day

    upkeep handled for you

  4. 04

    Enrichment

    Raw permit records provide limited context. Making them useful usually requires linking them to contractors, parcels, and related datasets, then enriching them with information such as contractor history, inspection outcomes, and parcel identifiers. That enrichment takes time and effort to build in-house. A commercial data provider usually offers enrichment out of the box.

  5. 05

    Delivery

    Access methods determine how readily the data reaches a team's workflow. A commercial provider typically offers an API, a web interface, and warehouse delivery, whereas an in-house pipeline supports only the formats the team builds. Shovels, for example, offers three delivery methods: a web app, API, or a custom data feed for enterprise customers.

  6. 06

    Total cost of ownership

    Because permits are public, the source data is free, but the full cost of an in-house approach includes engineering to build the pipeline, ongoing maintenance as sources change, and analyst time to clean and reconcile the output. A subscription consolidates those into a more predictable cost. Ultimately, balancing the cost depends on how many jurisdictions and how much enrichment a team requires.

How Shovels fits. Shovels applies the cleaning, standardization, and enrichment described above to permit, contractor, and property data from thousands of municipalities. Data delivery is available through the Shovels Online web application, the Shovels API, and Shovels Enterprise.

When does scraping permit data yourself make sense?

A DIY approach can absolutely make sense in certain situations. If you only need data from a single city or metro area, if you're doing a one-time analysis rather than building an ongoing data pipeline, or if you have engineering resources available and a well-defined set of requirements, collecting and managing the data yourself may be the more cost-effective option.

Shovels is designed for teams that need consistent coverage across many jurisdictions without having to build and maintain the underlying infrastructure. The trade-offs are worth understanding here as well.

Permit data is refreshed twice monthly, coverage is strongest in major metro areas (with some gaps in rural regions), and permits are not a perfect measure of completed construction since some projects are delayed, canceled, or never require a permit in the first place.

How should I evaluate buying vs. building?

Ultimately, the decision comes down to whether you want to own the permit data pipeline or simply use the data.

If you're evaluating a DIY approach, look beyond the cost of collecting records. Account for the engineering time required to build and maintain scrapers, the analyst hours needed to clean and standardize the data, and the effort involved in adding enrichment and keeping everything current as jurisdictions change their systems.

Once those costs are fully loaded, compare them against a Shovels plan covering the same markets. For some teams, especially those focused on a small number of jurisdictions, building in-house is likely the right choice. For teams that need broad coverage, consistent data quality, and minimal maintenance overhead, a managed dataset is the better investment.

This is especially the case for enterprise teams, who often get more value from receiving permit data directly in their warehouse (Snowflake, BigQuery, or Databricks) and integrating it into existing workflows.

Finally, it is worth asking what your team is really for. If your edge comes from making decisions with good data, that is where your people should spend their time, not on manually collecting and maintaining it.

Have more questions about build vs buy? Explore the Shovels permit database for free, or talk to us directly about your needs.

Want to see the difference on your own market?

Sign up free, pull the permits and contractors in a market you already know, and judge the coverage, deduplication, and enrichment for yourself.

Frequently asked questions

Why pay for Shovels instead of hiring an offshore team to scrape city websites?

A scraping team can pull permits, but you still own the cleaning, deduplication, standardization, and constant maintenance as cities change their sites, plus building any contractor or parcel enrichment yourself. Shovels delivers that as standardized, enriched data across 2,770+ jurisdictions, so you pay for usable intelligence rather than a pipeline you have to keep running. For offline jurisdictions, an offshore scraping team simply does not have the connections with local agencies needed to obtain records that never get published online. Shovels addresses this problem by building relationships with jurisdictions directly, ensuring we can access the valuable data you need.

What is the difference between a basic permit data dump and a permit intelligence platform?

A data dump is raw records in each city's own format. A permit intelligence platform standardizes those records into consistent fields and categories, deduplicates them, tags permit types, and links them to contractors and parcels, so the data is ready to analyze instead of ready to clean.

Can I just pull permits directly from city websites?

Yes, and for one city it can work. But even finding the right source is harder than it sounds. Is permitting handled by the city or the county? Does it depend on the permit type? Tracking down the correct portal for each jurisdiction is its own task before you write a line of scraping code. The difficulty scales with jurisdictions: every city has a different format and update method, and keeping many sources current and consistent is the part that turns a quick pull into an ongoing engineering project.

How does Shovels handle messy or incomplete permit records?

Shovels uses AI to clean and normalize inconsistent municipal text into standardized fields, deduplicate filings, and infer attributes like permit category. Field completeness still varies by jurisdiction, which is why Shovels publishes coverage detail rather than implying every record is complete.

Does Shovels have an API?

Yes. Shovels offers a REST API to pull permit data by geography and permit type, a web app for non-technical users, and direct warehouse sharing into Snowflake, BigQuery, and Databricks.

How is this different from pulling permits and adding parcel data myself?

You can join permits to parcels yourself, but it means sourcing and maintaining both datasets and reconciling addresses across them. Shovels links permits to standardized parcel IDs and property data out of the box, so the join is already done.