IDCite

A Large-Scale Multidisciplinary Citation Intent Dataset for Scholarly Knowledge Discovery

IDCite is a large-scale resource for citation-level analysis across multiple scientific disciplines. Seed publications were selected through a journal-stratified procedure covering 21 Essential Science Indicators (ESI) fields, 21 representative Web of Science categories, and 105 Q1 journals. The release contains 1,857,503 citation records linking 1,467,045 citing publications to 23,479 seed papers, together with citation contexts, intent labels, publication metadata, normalized entities, and an ontology-ready knowledge graph of 3,418,433 nodes and 6,855,117 edges.

Dataset Summary

A few headline numbers describe the scale of IDCite.

1,857,503
Citation events
1,467,045
Citing papers
23,479
Seed (cited) papers
21
ESI fields
105
Q1 journals sampled
Nov 2025
Data snapshot

Construction pipeline

From multidisciplinary seed selection to normalized, ready-to-use citation-event tables.

Figure 1. Overview of the IDCite construction workflow, linking input sources, the processing pipeline, and released records.
Scopus metadata WoS / 2024 JCR journal selection OpenAlex citation links Semantic Scholar citation contexts SynIntent intent classifier (GNN)
  1. 1

    Journal-stratified sampling

    21 ESI fields → 21 representative WoS categories → 5 Q1 journals each (105 journals).

  2. 2

    Seed-paper selection

    Top 5% most-cited papers retained independently within each journal → 23,479 seed papers.

  3. 3

    Citation harvesting

    OpenAlex citation links connect 1.47M citing papers to the seed corpus.

  4. 4

    Citation-event construction

    Each citing → seed link becomes an event carrying its citation context.

  5. 5

    Semantic enrichment

    SynIntent (GNN-based) assigns citation intents as weak semantic supervision.

  6. 6

    Entity normalization

    DOIs, authors, affiliations, journals, fields, and intents are normalized.

Citation intent distribution

Seven canonical intents after decomposing composite labels. The distribution preserves the natural imbalance of real scholarly citations — background dominates.

Canonical intentCount%
background1,660,89988.20
uses131,8277.00
similarities45,7532.43
motivation21,5551.14
differences15,6170.83
future work4,6540.25
extends2,8130.15

Intent labels are model-generated (SynIntent, Strict Acc. 0.67 / Weak Acc. 0.80 on MultiCite) and released as weak supervision — not manually curated gold labels.

Seed papers across 21 ESI fields

Journal-stratified sampling preserves disciplinary heterogeneity; field counts are not resampled to be balanced.

ESI field → representative WoS category (5 Q1 journals each)
ESI fieldWoS categoryJournals

Released files

Apache Parquet tables covering citation events, publication metadata, and normalized scholarly entities.

FileRowsColsDescription

Access & citation

Cite this dataset

@dataset{idcite2026,
  title     = {IDCite: A Large-Scale Multidisciplinary Citation Intent Dataset for Scholarly Knowledge Discovery},
  author    = {Nam, Seohyun},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.20796923}
}

Release lineage

  • Concept DOI (all versions) — 10.5281/zenodo.18410049
  • v1 · MDCite — 10.5281/zenodo.18410050
  • v2 · MDContextCite / EdgeCite — 10.5281/zenodo.18536895
  • v3 · IDCite — 10.5281/zenodo.20796923
PDF

Project & Dataset Documentation

The complete Version 3 reference — dataset design, release lineage, data sources, construction pipeline, schema definitions, sampling strategy, intent annotation, normalized entities, reproducibility, and responsible-use guidance.

Open documentation (PDF) →