Bulk Data Delivery
Fresh company intelligence, delivered to your cloud.
Complete company profiles covering details, headquarters, funding, investors, founders, headcount, and firmographics, delivered securely to Amazon S3, Snowflake, or Google BigQuery on the schedule you choose.
01 — Destinations
Deliver data where you need it
Data arrives in place, on the platform your team already runs on. No manual imports, no brittle file transfers. Apache Parquet is recommended for S3: faster queries at a fraction of the storage. Snowflake and BigQuery are shared natively and queryable the moment they land.
-
Amazon S3
Delivered straight to your S3 bucket via secure cross-account access or scoped bucket credentials.
Apache Parquet · recommended CSV -
Snowflake
Access data in place through Snowflake Secure Data Sharing. No imports, no file transfers.
Secure Data Sharing -
Google BigQuery
Receive a shared BigQuery dataset and query immediately with Google's native analytics.
Shared Dataset
02 — Delivery model
Complete company snapshots
Full Snapshot
Every scheduled delivery contains the complete company dataset: a consistent, comprehensive copy of the entire universe, ideal for rebuilds and full refreshes. Each delivery fully replaces the last, so there are no diffs to reconcile and no incremental state to manage.
Best for a complete, consistent copy of the data
- 2026-09-01 Complete dataset latest
- 2026-08-01 Complete dataset replaced
- 2026-07-01 Complete dataset replaced
03 — Schedule
Automated cadence
- Weekly Keep fast-moving data current with the freshest possible cadence.
- Monthly A dependable rhythm for most analytics and enrichment workflows.
- Quarterly Periodic refreshes for research, modeling, and reporting.
04 — What's inside
Rich company profile data
-
companyCanonical name, domain, founding year, and description.
-
headquartersCity, region, and country of the primary HQ.
-
fundingTotal raised, round count, and the latest round's series, date, amount, and lead investor.
-
investorsKnown investors across the company's funding rounds.
-
foundersFounders and key leadership with their roles.
-
employee_countLatest estimated headcount.
-
firmographicsIndustry, business model, and offering attributes.
-
job_openingsOpen roles from the company's own job board, refreshed weekly: total count plus the departments, countries, and cities it is hiring into. Covers about 20,000 companies, so it is a signal on the subset that publishes a board, not a whole-corpus field.
05 — Use cases
Why teams take the whole dataset
Bulk delivery is the right shape when the job needs the whole population rather than an answer to one question — a denominator, a negative class, or a local join. Each guide covers one of those jobs in full.
-
AI training and grounding
Feed a model the whole population, including the companies no query would surface.
-
CRM enrichment at scale
Turn enrichment into a warehouse join instead of a per-record queue.
-
Entity resolution
Build a master company table with one stable identity per company.
-
TAM and market sizing
Count segments against the full population instead of extrapolating from a sample.
-
In-product features
Serve autocomplete and company profiles from your own infrastructure, not a vendor's.
-
Lead scoring models
Train on a real negative class, then score the whole market in one pass.
06 — Why Canonical
Why Canonical
-
Secure by design
Cloud-native sharing through cross-account S3 access, Snowflake shares, and BigQuery datasets. Data stays in platforms you trust.
-
Reliable schedules
Automated weekly, monthly, or quarterly deliveries your pipelines can depend on.
-
Complete snapshots
A consistent, comprehensive copy of the full company dataset with every scheduled delivery.
-
Native integrations
Deliveries built on each platform's own sharing mechanisms: Amazon S3, Snowflake, and Google BigQuery.
-
Trusted intelligence
High-quality, canonical company profiles built for products, models, and research.
07 — FAQ
Frequently asked questions
Which cloud platforms are supported for data delivery?
Amazon S3, Snowflake, and Google BigQuery. S3 delivers Apache Parquet (recommended) or CSV. Snowflake and BigQuery use native in-place sharing, so data is queryable immediately.
How often is the data delivered?
On an automated schedule that fits your workflow: weekly, monthly, or quarterly. Every delivery is automated, secure, and designed to arrive reliably and on time.
What format is the data delivered in?
Apache Parquet (recommended) or CSV for Amazon S3. Snowflake and BigQuery deliveries are queryable tables via secure sharing, so no file format is required.
How is delivery secured?
Through cloud-native mechanisms: scoped cross-account access or bucket credentials for S3, Snowflake Secure Data Sharing, and shared BigQuery datasets. Delivery runs entirely on the cloud providers' own sharing infrastructure, with no third-party intermediaries or manual file transfers.
What company attributes are included?
Company details, headquarters and location, funding history, investors, founders, employee count, and industry and firmographic attributes.
Can I get a sample dataset before committing?
Yes. Book a demo and we'll walk you through a sample dataset.
Simplify your company data pipeline.
Whether you're building analytics, AI, research, or operational workflows, we'll help you find the right delivery option and send a sample.