ЭФФЕКТИВНАЯ КЛАССИФИКАЦИЯ НАУЧНЫХ ТЕКСТОВ С ИСПОЛЬЗОВАНИЕМ МОДЕЛИ SCIBERT И МЕТОДА НИЗКОРАНГОВОЙ АДАПТАЦИИ.
Ключевые слова:
NLP, SciBERT, BERT, LoRA, классификация научных текстов, машинное обучение, текстовая классификация, TF-IDF, логистическая регрессияАннотация
В данной работе предлагается эффективный метод классификации научных текстов на основе модели SciBERT, дополнительно обученной с использованием техники LoRA с целью снижения вычислительных затрат. В ходе исследования была рассмотрена задача классификации научных абстрактов, при этом использовались два набора данных: PubMed-RCT20k — для биомедицинских текстов и SciCite — для классификации целей цитирования. Производительность базовой модели на основе логистической регрессии и TF-IDF была сопоставлена с результатами улучшенной модели SciBERT + LoRA. Результаты исследования показали, что модель SciBERT + LoRA превосходит модель логистической регрессии по показателям точности (Accuracy), макро F1-мере (Macro F1 score) и функции потерь (Loss), при этом требует значительно меньших вычислительных ресурсов. Проведённые эксперименты доказали, что адаптация модели SciBERT с помощью LoRA является эффективным подходом в задачах классификации научных текстов, особенно в условиях ограниченных вычислительных ресурсов. Кроме того, метод LoRA позволяет достигать высокой производительности без дорогостоящей вычислительной инфраструктуры, что делает его оптимальным решением для исследовательских групп с ограниченными вычислительными возможностями.
Библиографические ссылки
. Jamshidi, S., Mohammadi, M., Bagheri, S., Esmaeili Najafabadi, H., Rezvanian, A., Gheisari, M., Ghaderzadeh, M., Shahabi, A. S., & Wu, Z. (2024). Effective text classification using BERT, MTM LSTM, and DT. Data & Knowledge Engineering, 151, 102306. https://doi.org/10.1016/j.datak.2024.102306.
. Zaman‑Khan, H., et al. (2024). Enhancing text classification using BERT: A transfer‑learning approach. Journal Name, Volume(Issue), Page range. https://doi.org/10.13053/cys-28-4-5290
. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv. https://doi.org/10.48550/arXiv.1810.04805
. Li, J., Wei, Q., Ghiasvand, O. et al. A comparative study of pre-trained language models for named entity recognition in clinical trial eligibility criteria from multiple corpora. BMC Med Inform Decis Mak 22 (Suppl 3), 235 (2022). https://doi.org/10.1186/s12911-022-01967-7
. Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, Jaewoo Kang, BioBERT: a pre-trained biomedical language representation model for biomedical text mining, Bioinformatics, Volume 36, Issue 4, February 2020, Pages 1234–1240, https://doi.org/10.1093/bioinformatics/btz682
. Jamshidi, S., Mohammadi, M., Bagheri, S., Esmaeili Najafabadi, H., Rezvanian, A., Gheisari, M., Ghaderzadeh, M., Shahabi, A. S., & Wu, Z. (2024). Effective text classification using BERT, MTM LSTM, and DT. Data & Knowledge Engineering, 151, 102306. https://doi.org/10.1016/j.datak.2024.102306
. Fatwanto, A., Zamakhsyari, F., Ndungi, R., & Fitriyani, N. L. (2024). A systematic literature review of BERT‑based models for natural language processing tasks. Infotel, 16(3), 1206. https://doi.org/10.20895/INFOTEL.V16I3.1206
. Zangari, A.; Marcuzzo, M.; Rizzo, M.; Giudice, L.; Albarelli, A.; Gasparetto, A. Hierarchical Text Classification and Its Foundations: A Review of Current Research. Electronics 2024, 13, 1199. https://doi.org/10.3390/electronics13071199
. Hu, E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., & Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv. https://doi.org/10.48550/arXiv.2106.09685
. Mao, Y., Ge, Y., Fan, Y., Xu, W., Mi, Y., Hu, Z., & Gao, Y. (2024). A survey on LoRA of large language models. arXiv. https://doi.org/10.48550/arXiv.2407.11046
. Dernoncourt, F., & Lee, J. Y. (2017). PubMed 200k RCT: a dataset for sequential sentence classification in medical abstracts. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers) (pp. 308–313). Asian Federation of Natural Language Processing. https://aclanthology.org/I17-2052/
. Cohan, A., Ammar, W., van Zuylen, M., & Cady, F. (2019). Structural Scaffolds for Citation Intent Classification in Scientific Publications [Dataset and model description]. Proceedings of NAACL‑HLT 2019. https://aclanthology.org/N19-1361/
. Le, T.-D., Nguyen, T. Ti, Ha, V. Nguyen, Chatzinotas, S., Jouvet, P., & Noumeir, R. (2025). The impact of LoRA adapters on LLMs for clinical text classification under computational and data constraints. IEEE Access, 13, 109365-109377. https://doi.org/10.1109/ACCESS.2025.3582037
. Hu, J., Liao, X., Gao, J., Qi, Z., Zheng, H., & Wang, C. (2024). Optimizing large language models with an enhanced LoRA fine-tuning algorithm for efficiency and robustness in NLP tasks. In 2024 4th International Conference on Communication Technology and Information Technology (ICCTIT) (pp. 526-530). Guangzhou, China. https://doi.org/10.1109/ICCTIT64404.2024.10928552
. He, S., Lei, Y., Zhang, Y., Xie, K., & Sharma, P. K. (2023). Parameter-efficient log anomaly detection based on pre-training model and LoRA. In Proceedings of the 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE), Florence, Italy, 207–217. https://doi.org/10.1109/ISSRE59848.2023.00038
. Frisoni, G., Moro, G., Carlassare, G., & Carbonaro, A. (2022). Unsupervised Event Graph Representation and Similarity Learning on Biomedical Literature. Sensors, 22(1), 3. https://doi.org/10.3390/s22010003
. Wang, X., Aitchison, L., & Rudolph, M. (2023). LoRA ensembles for large language model fine-tuning. arXiv. https://doi.org/10.48550/arXiv.2310.00035
. Muralidharan, B., Beadles, H., Marzban, R., & Mupparaju, K. S. (2024). Knowledge AI: Fine-tuning NLP models for facilitating scientific knowledge extraction and understanding. arXiv. https://doi.org/10.48550/arXiv.2408.04651
. Satya S Sahoo, Joseph M Plasek, Hua Xu, Özlem Uzuner, Trevor Cohen, Meliha Yetisgen, Hongfang Liu, Stéphane Meystre, Yanshan Wang, Large language models for biomedicine: foundations, opportunities, challenges, and best practices, Journal of the American Medical Informatics Association, Volume 31, Issue 9, September 2024, Pages 2114–2124, https://doi.org/10.1093/jamia/ocae074
. Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., & Poon, H. (2021). Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 3(1), Article 2. https://doi.org/10.1145/3458754
. Maheshwari, H., Singh, B., & Varma, V. (2021). SciBERT sentence representation for citation context classification. In Proceedings of the Second Workshop on Scholarly Document Processing (pp. 130–133). Association for Computational Linguistics. https://aclanthology.org/2021.sdp-1.17/
. Li, T., Wang, J., Zhang, Y., Li, S., & Chen, L. (2025). Adapting pretrained language models for citation classification via self-supervised contrastive learning. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (pp. 1541–1552). Association for Computing Machinery. https://doi.org/10.1145/3711896.3736829
. Priya, B. R., & Shettar, R. (2025). Biomedical text classification using transformer models for early autism prediction. TechRxiv. https://doi.org/10.36227/techrxiv.175623876.62294534/v1
. Gu, N., & Hahnloser, R. (2024). Controllable citation sentence generation with language models. In Proceedings of the Fourth Workshop on Scholarly Document Processing (SDP 2024) (pp. 22–37). Association for Computational Linguistics. https://aclanthology.org/2024.sdp-1.4/
Опубликован
Как цитировать
Выпуск
Раздел
Категории
Лицензия
Copyright (c) 2025 Динара Касымова, Айнур Турсынхан , Мұса Тұрдалыұлы , Айгерим Еримбетова , Нуржан Мукажанов

Это произведение доступно по лицензии Creative Commons «Attribution-NonCommercial-NoDerivatives» («Атрибуция — Некоммерческое использование — Без производных произведений») 4.0 Всемирная.











