R&D intelligence · open access

Spot the next breakthrough 12–18 months before the journals.

We mirror every public preprint server — arXiv, bioRxiv, medRxiv, Crossref, OpenAlex, plus lab & corporate research blogs — and score each paper for novelty with a transparent heuristic. No paywalls. No AI hallucination. Just signal.

Updated live each run 0 paid APIs 6 open sources
Why this matters

Four signals, one dashboard

Each card is computed from public preprints — no API keys, no proprietary data. Hover any KPI card on the live dashboard for definitions.

{Total papers scraped}

1449

Across arXiv, bioRxiv & medRxiv, OpenAlex, Crossref, and lab RSS feeds. De-duplicated before counting.

Browse the corpus

{Top emerging category}

Artificial Intelligence

215 papers in the lead subject — the live trend-curve for 2025–2026.

See the distribution
N

{Highest signal item}

The Causal Artificial Intelligence Clinician for early haemodynamic management of septic shock in ICU

Score: 91 / 100

The single most novel-feeling paper in the corpus. Open the table to read the abstract & PDF.

Open the table

{Entities monitored}

5090

De-duplicated authors, labs & institutions in the corpus. Breadth of contributors — a leading indicator of field momentum.

Drill into the data
Methodology & how to read

Everything on this page is instructive by design

We do not want you to just see data — we want you to understand what it means, where it came from, and what to do with it.

How this dashboard works

A live mirror of the public preprint ecosystem, scored for novelty with a transparent heuristic. Every paper below was harvested in the last run from free, public sources.

What is a "preprint"?

A manuscript posted by researchers to an open-access server before or instead of formal peer review. Preprints reveal early-stage work before it appears in journals — 12–18 months earlier than commercial / journal channels.

Signal score formula

score = min(len(abstract)/50, 60)
+ min(num_keywords * 4, 25)
+ min(num_novel_keywords * 4, 15)
+ min(max(num_authors-1, 0) * 0.8, 6)

Tier thresholds

High ≥70   Medium ≥45   Watchlist <45

Sources & what they cover

  • arXiv — physics, math, CS, q-bio. 30 categories scraped at 40 papers each (2 pages).
  • bioRxiv & medRxiv — biological & medical preprints. 60-day windowed JSON API.
  • OpenAlex (preprint) — open scholarly graph. Cross-domain preprint discovery (ChemRxiv is Cloudflare-gated).
  • Crossref (posted-content) — publisher-neutral index of preprints & conference pre-prints.
  • Lab & corporate blogs — DeepMind, OpenAI, Google Research, Hugging Face, AWS ML, IBM Research, MIT Tech Review, Quanta, Phys.org, Ars Technica, ScienceDaily, Nature, MIT News, Stanford HAI.

Constraints

  • Zero paid API dependencies — everything is open.
  • Randomized UA, 1.5–4.5s polite delay, exponential backoff on 429/403.
  • No AI/LLM calls in scoring; signal score is a transparent heuristic.
  • Built autonomously, single-file dashboard, zero build step.

How to read this dashboard

  1. Read the KPIs top-down. Total papers → top category → highest signal item.
  2. Cross-check the doughnut. If the bulk is concentrated in one subject, that is a noise floor — look for smaller slices.
  3. Read the keyword chart. Rapid co-occurrence is a proxy for "what are multiple labs writing about".
  4. Filter the table. Try "High Priority" first; use search for exact terms (e.g. "mamba", "MXene", "diffusion transformer").
  5. Export CSV for offline triage or to feed into your own LLM-augmented analysis pipeline.

Watch for

Rapid jumps in any single keyword (new subfield), emergence of cross-domain authors (e.g. materials + ML), and any paper that scores ≥ 90 — the rare 1-in-100 candidate worth deep-reading.

Glossary & what to look for

preprint — author-deposited manuscript not yet peer-reviewed. Earliest public signal of new research.

arXiv category codes — e.g. cs.AI = Computer Science · Artificial Intelligence.

cond-mat.* — condensed-matter physics: mtrl-sci materials, mes-hall mesoscale, soft soft matter, str-el strongly correlated, supr-con superconductivity.

q-bio.* — quantitative biology: BM biomolecules, NC neurons, QM quant bio, GN genomics, SC synthetic biology.

eess.* — electrical engineering: AS audio/speech, SP signal processing, IV image/video.

Signal score — composite of abstract length, keyword density, novel-term count, and author count. Higher = more novel-feeling.

Tier — High (≥70) / Medium (≥45) / Watchlist (<45). Triage aid, not a verdict.

"Other" bucket — tail of small categories rolled together so the chart stays readable.

Interactive research discovery

Every paper in the corpus

Search, filter, sort, export. Click any column header to sort.

Showing 0 rows
Title & link Authors Category Keywords Date Score

Row color band on the left = tier: High   Medium   Watchlist

Why it is different

Traditional R&D intel vs. R&D Signal

A direct comparison of how patent attorneys, VC technical partners, and corp R&D executives usually track trends — vs. how this dashboard does it.

Traditional approach

Journal & patent watch

  • Reads journals 12–24 months after preprint posting.
  • Pays for $50k/yr Elsevier / Clarivate / PatSnap subscriptions.
  • Manual keyword tracking & analyst hours.
  • Misses cross-domain & non-patent disclosures.
  • Mostly retrospective — tells you what happened, not what is forming.
Traditional approach

Conference crawl

  • Manually tracks 5–20 venues (NeurIPS, ICLR, ASCO, MRS…).
  • After acceptance — typically 6–12 months post-submission.
  • Heavily skewed toward a single field per track.
  • Easy to miss cross-domain breakthroughs.
  • No quantitative novelty signal — all manual.
This dashboard

R&D Signal — autonomous

  • Live mirror of 6 open-access sources — no login required.
  • Preprints & posted-content — weeks-to-months ahead of journals.
  • Heuristic novelty score (1–100) per paper, transparent formula.
  • Cross-domain by construction (CS · cond-mat · q-bio · eess · ...).
  • Zero subscription cost, zero AI hallucination, runs in seconds.
  • Single-file HTML you can host anywhere — even offline.

Built on data from the world's leading open-access preprint servers

arXiv bioRxiv medRxiv Crossref OpenAlex Open Access