- Data Descriptor
- Open access
- Published:
- Muhammad Alfian1,
- Daniel Siahaan1,
- Harum Munazharoh2,
- Edio da Costa3 &
- …
- Eric Pardede4
Scientific Data (2026) Cite this article
We are providing an unedited version of this manuscript to give early access to its findings. Before final publication, the manuscript will undergo further editing. Please note there may be errors present which affect the content, and all legal disclaimers apply.
Abstract
Part-of-speech (POS) tagging is an essential process that can significantly impact other text analysis tasks. Indonesian POS tagging has currently achieved a high accuracy rate using the BiLSTM method. However, this value is still not optimal due to challenges in identifying word classes for words with multiple meanings (ambiguity) and out-of-vocabulary (OOV) words. One reason is the lack of a consistent and verifiable corpus that can serve as a reliable, verifiable data source. Therefore, this study proposes a mechanism to develop a corpus by adapting existing corpora from Fu et al., which consist of 21,024 sentences, and enhancing them with linguist-based label validation. The corpus development process involves six steps: linguist selection, inconsistency identification, preprocessing, inconsistency detection, inconsistency correction, and corpus evaluation. The evaluation was conducted on a subset of 2,457 validated single-clause sentences, comparing the initial filtered corpus with IDeal, the proposed corpus. Results show that the number of label variations in the proposed corpus has decreased compared to the initial filtered corpus, as indicated by a reduction in entropy from 0.04 to 0.02. The evaluation demonstrates that IDeal outperforms previous benchmarks. Furthermore, within this restricted single-clause context, the IDeal corpus improved model performance in general and ambiguous cases. The model’s Macro F1 Score increased by approximately 8% for general cases and about 9% for ambiguous cases. The dataset is available to download from https://doi.org/10.5281/ZENODO.20159181.
Acknowledgements
This research is funded by the Indonesian Endowment Fund for Education (LPDP) on behalf of the Indonesian Ministry of Higher Education, Science and Technology and managed under the EQUITY Program (Contract No 4299/B3/DT.03.08/2025 & No 3029/PKS/ITS/2025). We would like to express our gratitude to Ade Putri Riajang Pandan Tunjung Sari, Afidati Lelani Putri, Alvina Yustiyaningrum, Arina Nida Arrusyda, Arizqa Novi Ramadhani, Arynda Natasha Dewanty, Aulia Destya Putri, Azlin Huwaida Azzahra, Betty Khasandra Pujayanti, Bilqis Tirtakayana, Dewi Ayu Permatasari, Edgar Allan Stefan, Erika Asni Sabrina, Fadhila Khusnul Na’imah, Fitri Aljazera, Haikal Wiranata, Henik Fuji Rahmawati, Hilmy Hidayana, Muhimatul Khoiriyah, Nailul Hikmah, Nurkumala Dewi, Talita Hariyanto, Theophillus Bagas Adhiatma, Ulfa Nadhiroh Mukminin, Vista Artamarista Hapsari Kusuma, Zahra Putri Pratiwig, Zhafira Afra Putri Hermawan for their contributions in annotating our data.
Ethics declarations
Competing interests
The authors declare no competing interests.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Reprints and permissions
About this article
Cite this article
Yuhana, U.L., Alfian, M., Siahaan, D. et al. A Part-of-Speech Tagging Corpus for Bahasa Indonesia with Linguist-Based Label Validation. Sci Data (2026). https://doi.org/10.1038/s41597-026-08046-w
Download citation
Received:
Accepted:
Published:
DOI: https://doi.org/10.1038/s41597-026-08046-w