A Part-of-Speech Tagging Corpus for Bahasa Indonesia with Linguist-Based Label Validation

Nature作者:Umi Laili Yuhana2026年8月12日正文已收录本站
  • Data Descriptor
  • Open access
  • Published:
  • Muhammad Alfian1,
  • Daniel Siahaan1,
  • Harum Munazharoh2,
  • Edio da Costa3 &
  • …
  • Eric Pardede4 

Scientific Data (2026) Cite this article

We are providing an unedited version of this manuscript to give early access to its findings. Before final publication, the manuscript will undergo further editing. Please note there may be errors present which affect the content, and all legal disclaimers apply.

Abstract

Part-of-speech (POS) tagging is an essential process that can significantly impact other text analysis tasks. Indonesian POS tagging has currently achieved a high accuracy rate using the BiLSTM method. However, this value is still not optimal due to challenges in identifying word classes for words with multiple meanings (ambiguity) and out-of-vocabulary (OOV) words. One reason is the lack of a consistent and verifiable corpus that can serve as a reliable, verifiable data source. Therefore, this study proposes a mechanism to develop a corpus by adapting existing corpora from Fu et al., which consist of 21,024 sentences, and enhancing them with linguist-based label validation. The corpus development process involves six steps: linguist selection, inconsistency identification, preprocessing, inconsistency detection, inconsistency correction, and corpus evaluation. The evaluation was conducted on a subset of 2,457 validated single-clause sentences, comparing the initial filtered corpus with IDeal, the proposed corpus. Results show that the number of label variations in the proposed corpus has decreased compared to the initial filtered corpus, as indicated by a reduction in entropy from 0.04 to 0.02. The evaluation demonstrates that IDeal outperforms previous benchmarks. Furthermore, within this restricted single-clause context, the IDeal corpus improved model performance in general and ambiguous cases. The model’s Macro F1 Score increased by approximately 8% for general cases and about 9% for ambiguous cases. The dataset is available to download from https://doi.org/10.5281/ZENODO.20159181.

Acknowledgements

This research is funded by the Indonesian Endowment Fund for Education (LPDP) on behalf of the Indonesian Ministry of Higher Education, Science and Technology and managed under the EQUITY Program (Contract No 4299/B3/DT.03.08/2025 & No 3029/PKS/ITS/2025). We would like to express our gratitude to Ade Putri Riajang Pandan Tunjung Sari, Afidati Lelani Putri, Alvina Yustiyaningrum, Arina Nida Arrusyda, Arizqa Novi Ramadhani, Arynda Natasha Dewanty, Aulia Destya Putri, Azlin Huwaida Azzahra, Betty Khasandra Pujayanti, Bilqis Tirtakayana, Dewi Ayu Permatasari, Edgar Allan Stefan, Erika Asni Sabrina, Fadhila Khusnul Na’imah, Fitri Aljazera, Haikal Wiranata, Henik Fuji Rahmawati, Hilmy Hidayana, Muhimatul Khoiriyah, Nailul Hikmah, Nurkumala Dewi, Talita Hariyanto, Theophillus Bagas Adhiatma, Ulfa Nadhiroh Mukminin, Vista Artamarista Hapsari Kusuma, Zahra Putri Pratiwig, Zhafira Afra Putri Hermawan for their contributions in annotating our data.

Author information

Authors and Affiliations

  1. Department of Informatics, Institut Teknologi Sepuluh Nopember, Surabaya, Indonesia

    Umi Laili Yuhana, Muhammad Alfian & Daniel Siahaan

  2. Department of Indonesian Language and Literature, Universitas Airlangga, Surabaya, Indonesia

    Harum Munazharoh

  3. School of Engineering and Science, Dili Institute of Technology, Dili, Timor-Leste

    Edio da Costa

  4. School of Computing, Engineering and Mathematical Sciences, La Trobe University, Melbourne, Australia

    Eric Pardede

Authors

  1. Umi Laili Yuhana
  2. Muhammad Alfian
  3. Daniel Siahaan
  4. Harum Munazharoh
  5. Edio da Costa
  6. Eric Pardede

Corresponding author

Correspondence to Umi Laili Yuhana.

Ethics declarations

Competing interests

The authors declare no competing interests.

Additional information

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Yuhana, U.L., Alfian, M., Siahaan, D. et al. A Part-of-Speech Tagging Corpus for Bahasa Indonesia with Linguist-Based Label Validation. Sci Data (2026). https://doi.org/10.1038/s41597-026-08046-w

Download citation

  • Received:

  • Accepted:

  • Published:

  • DOI: https://doi.org/10.1038/s41597-026-08046-w