- Article
- Open access
- Published:
npj Computational Materials (2026) Cite this article
We are providing an unedited version of this manuscript to give early access to its findings. Before final publication, the manuscript will undergo further editing. Please note there may be errors present which affect the content, and all legal disclaimers apply.
Abstract
Membrane-based separation is a critical technology for energy and environmental processes, positioning it at the forefront of materials science research. To facilitate the data-driven advances in this field, we developed an integrated natural language processing pipeline including a domain-tailored transformer model for literature mining. One major challenge is data imbalance which biases the model toward majority classes and leads to overfitting on dominant patterns. We introduced two complementary strategies to direct the model’s attention toward minority classes: (1) a data augmentation method that prevents data leakage, and (2) model fine-tuning using a customized loss function that integrates class-weighted cross-entropy and focal loss. The resulting model MembraneBERT achieves an average Precision of 0.906, Recall of 0.796, and F1-score of 0.844, respectively. Applied to 3,580 additional articles, the model extracted 2,283 structured entries on membranes and their gas separation performance, offering a practical framework for AI-driven extraction of membrane performance data to support subsequent data-driven research.
Subjects
Acknowledgements
This work was supported by the National Natural Science Foundation of China (Grant no. 22408071), the Hainan Provincial Natural Science Foundation of China (Grant no. 626MS0099 and 525QN256), and the Hainan Province Science and Technology Special Fund (Grant no. ZDYF2025SHFZ025). The funder played no role in study design, data collection, analysis and interpretation of data, or the writing of this manuscript.
Ethics declarations
Competing interests
The authors declare no competing interests.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Reprints and permissions
About this article
Cite this article
Cao, Y., Jiang, Y., Yang, Y. et al. MembraneBERT: a domain-tailored NLP model for membrane-based gas separation data extraction. npj Comput Mater (2026). https://doi.org/10.1038/s41524-026-02227-2
Download citation
Received:
Accepted:
Published:
DOI: https://doi.org/10.1038/s41524-026-02227-2