News & Sentiment

From AltData.wiki, The Alternative Data Encyclopedia

News & Sentiment is one of the categories of alternative data covered by The Alternative Data Encyclopedia: NLP scoring of news, press releases, filings, earnings-call transcripts and social posts into machine-readable tone, novelty and event signals per company. The category converts unstructured text flow into quantitative inputs that investment models can consume at scale.

RavenPack began scoring financial news for funds in the early 2000s, and Paul Tetlock's 2007 Journal of Finance study linking media pessimism to subsequent returns provided the academic anchor for the field. MarketPsych has produced point-in-time financial sentiment data since 2004 and co-authored much of the underlying research. The category has since extended from equities to credit, FX, commodities and crypto, and now often ships as annotation layers for training large language models.

The signal

NLP scoring of news, press releases, filings, earnings-call transcripts and social posts into machine-readable tone, novelty and event signals per company. The category converts unstructured text flow into quantitative inputs that investment models can consume at scale.

Article-level records carry entity tags, relevance scores, sentiment scores and event labels (guidance changes, litigation, management departures), together with source metadata and links. Aggregation layers produce company-day sentiment, momentum and novelty series plus cross-sectional dispersion measures. Vendors increasingly ship factor-ready outputs such as controversy indices and earnings-call tone deltas.

Why investors pay for it

Information reaches prices with a lag, so shifts in media tone and story novelty anticipate short-horizon returns, volatility spikes and reversal windows. Funds use sentiment momentum around earnings, controversy screens as risk overlays, and transcript-tone deterioration to flag weakening fundamentals before estimate revisions. Story-level deduplication ensures a single scoop is counted once rather than once per syndicated copy.

Pipelines ingest licensed newswires, regulatory feeds, transcript services and social APIs, then tag entities using curated dictionaries with disambiguation by ticker and corporate hierarchy. Language models tuned for financial tone score each item; established vendors publish dozens of sentiment-related fields per entity and cover a dozen or more languages. Point-in-time discipline is central: timestamps reflect publication time, and dictionaries and models are versioned so backtests remain reproducible.

Who uses it

Quantitative equity funds consume sentiment and novelty scores as model features; event-driven desks monitor controversy and novelty alerts around catalysts. Risk teams apply negative-news screening to portfolios and counterparty surveillance.

Questions to ask vendors in this category

Which sources are licensed versus scraped, and what happens to continuity when a license lapses? What is feed latency from publication to delivered score? How do you disambiguate entities with common names and complex hierarchies? Are historical scores recomputed when models change, or shipped point-in-time? How long is the clean backtest history across asset classes?

Complementary signals

This signal pairs naturally with adjacent categories of the encyclopedia:

Caveats and limitations

Sentiment scores are model-dependent and unstable around sarcasm, hedging and translated text. Coverage volume scales with market capitalization, so unnormalized activity proxies company size rather than information. Clustering errors double-count events, and silent vendor model upgrades can rewrite historical values.

Compliance and legal considerations

Redistributing article text requires publisher agreements, while derived scores are generally safer to ship. Social-media ingestion adds platform terms-of-service constraints, and text-and-data-mining exceptions in EU copyright law do not override contractual restrictions.

Further reading

Providers in this category

The register lists 37 companies for this signal family:

Top premium datasets (5)

Top free datasets (5)