Job market data: datasets and providers
From AltData.wiki, The Alternative Data Encyclopedia | Primer | 3 min read
The datasets and providers behind labor-market signals: job postings, payroll, employee profiles and job-change data — what each source measures, its caveats and how to choose one.
Why labor data leads fundamentals
Hiring is a commitment of budget and management attention, which makes it a revealed-preference indicator of a company's expectations. Companies post roles when they expect demand and cut roles when they see trouble, usually weeks or months before the results show up in filings. Labor data therefore leads fundamentals in a way that quarterly numbers cannot, and it covers private companies that publish no filings at all.
The category splits into four complementary sources, each measuring a different layer of the labor market: job postings (intent to hire), payroll and benefits data (realized employment and compensation), employee profiles and job changes (who works where), and employer-review or skills data (sentiment and capability). The signal is strongest when these layers are combined — a company scaling sales postings while payroll headcount is flat is behaving differently from one doing both.
Job postings as the core signal
Job postings are the most widely used labor signal because they are public, near real time and granular. Vendors crawl company career pages and job boards, deduplicate listings, normalize titles and geographies, and timestamp when a role appears and disappears. The resulting series tracks hiring intent by function, seniority and location, which analysts translate into growth, expansion and technology-adoption signals.
The mechanics and failure modes are covered in Job postings as an investment signal. In short: deduplication quality, the split between new and reposted roles, and coverage of small firms all vary by vendor, and those differences explain most of the disagreement between competing datasets. The jobs-workforce category article indexes the full provider list.
Providers of job-posting data
LinkUp crawls company career sites directly rather than aggregating job boards, which gives cleaner deduplication and a truer picture of newly posted roles. Lightcast (formerly Emsi Burning Glass) and Coresignal process postings at scale, adding firmographic mapping and longitudinal history. Thinknum and JobsPikr offer programmatic access to postings for quants building their own pipelines.
The job boards themselves are also sources. Indeed publishes hiring-lab and job-posting data, and its listings flow into many Kaggle and marketplace datasets. Greenwich.HR and TalentNeuron layer compensation and talent analytics on top. The register's jobs-workforce providers list every option with delivery and price band.
Payroll, profiles and job-change data
Payroll data measures realized employment rather than intent. ADP publishes aggregated employment and pay series that economists treat as a high-frequency read on the labor market, and its anonymized workforce data feeds commercial datasets. LinkedIn contributes its Economic Graph of profiles, skills and transitions, used for workforce-flow and talent-migration research.
Job-change data captures movement between roles, a strong signal for growth and churn. Revelio Labs reconstructs workforce dynamics from public profiles and postings; People Data Labs and Live Data Technologies sell person-level and change-detection feeds. These sources fill the gap between a company's stated hiring intent and its actual headcount.
Datasets indexed in the register
Beyond the commercial vendors, the register indexes open and marketplace datasets that give a team a starting point without an enterprise contract. TheirStack job postings is the category's first-party benchmark; LinkUp's structured job listings and Revelio human-capital dynamics are the premium reference sets.
On the open side, Kaggle's Indeed job postings and the various LinkedIn/Indeed posting mirrors provide snapshots for prototyping and academic study, while Coresignal's open employee dataset offers a broad employee-profile base. These are listed alongside every other job dataset in the jobs-workforce category.
Caveats and how to choose
The main caveats are deduplication, coverage and point-in-time integrity. Postings are duplicated across boards and reposted by recruiters, so raw counts overstate hiring; different vendors deduplicate differently, so two series can disagree on the same company. Coverage skews toward online-friendly firms and away from small or manual-hiring employers. And historical postings files are often backfilled, so a backtest must use snapshots rather than a single downloaded history.
Choosing a provider comes down to the question being asked. For a broad, historically consistent series, LinkUp or Lightcast are the default; for person-level and transition detail, Revelio or People Data Labs; for a low-cost prototype, an open Kaggle or marketplace snapshot. The same evaluation discipline from How funds evaluate alternative data vendors applies, with special attention to the compliance story behind profile data, which can implicate privacy rules.
Frequently asked questions
- What is job market data in alternative data?
- Data describing the labor market — job postings, payroll and employment, employee profiles and job changes — used to forecast company growth, cost control and macro employment conditions.
- Which provider has the cleanest job-posting data?
- LinkUp is the common reference because it crawls company career sites directly and deduplicates carefully; Lightcast and Coresignal are comparable at scale. The best choice depends on history depth and entity mapping needs.
- Can I get job data for free?
- Yes for prototyping: Kaggle and marketplace snapshots (including Indeed postings) are free, and Coresignal publishes an open employee dataset. Production-quality series generally require a paid license.
- Does job data cover private companies?
- Yes — postings are public regardless of ownership, which makes labor data one of the few signals that covers private companies before any filing exists.
- Why do two job datasets disagree on the same company?
- Mostly deduplication and coverage. Vendors crawl different boards, deduplicate reposted roles differently, and cover different subsets of small employers, so raw counts and trends can differ.