Research Article | Open Access | Download PDF
Volume 13 | Issue 9 | Year 2026 | Article Id. IJECE-V13I9P110 | DOI : https://doi.org/10.14445/23488549/IJECE-V13I9P110A Context-Aware, Feature-Enriched Bidirectional Long Short-Term Memory Framework for Automatic Document Readability Assessment
Rajesh Kandakatla, V. Dhilip kumar
| Received | Revised | Accepted | Published |
|---|---|---|---|
| 30 Jul 2026 | 16 Sep 2026 | 18 Sep 2026 | 29 Sep 2026 |
Citation :
Rajesh Kandakatla, V. Dhilip kumar, "A Context-Aware, Feature-Enriched Bidirectional Long Short-Term Memory Framework for Automatic Document Readability Assessment," International Journal of Electronics and Communication Engineering, vol. 13, no. 9, pp. 160-180, 2026. Crossref, https://doi.org/10.14445/23488549/IJECE-V13I9P110
Abstract
Automatic readability assessment is important in the domain of educational content selection, language learning, digital accessibility, and personalized reading assistance. The traditional readability formulas quickly show a high dependence on bare-bones factors, and as such, they can not adequately assess the context of meaning, syntactic config, semantic entropy, discourse cohesion, and arranged structure types for degrees of ease. This work presents a framework for automatic document readability assessment which takes into consideration the context and other related features. They mixed several elements in their model, including RoBERTa-based contextual embeddings, a Bidirectional Long Short-Term Memory (Bi-LSTM) encoder, an attention mechanism, multidimensional linguistic features, gated feature fusion, and an ordinal-aware learning objective. These lexical, syntactic, semantic, and discourse cues are fused using an attention-weighted contextual representation to obtain a document-level feature vector. In the two-stage transfer-learning strategy, first, we pre-train our encoder on the CEFR-Based Sentence Profile corpus and then fine-tune it with the OneStopEnglish document-level corpus. The final objective involves the use of weighted categorical cross-entropy in conjunction with an ordinal-distance penalty to take into account class imbalance, as well as the seriousness of the error between more distant readability classes. Evaluation metrics: The evaluation metrics used include accuracy, macro-F1, Mean Absolute Error (MAE), Quadratic Weighted Kappa (QWK), adjacent level accuracy, and Spearman rank correlation. The approximation evaluation shows that after using the CEFR-SP pretrained model, our accuracy is 93.19%, macro-F1 is 92.96%, MAE is 0.078, and QWK is 0.956. Ablation studies also show that the model performance increases progressively through incorporating bidirectional encoding, attention, linguistic-feature fusion, auxiliary pretraining, and ordinal supervision. This framework forms a basis for effective context-sensitive and order-aware document readability prediction.
Keywords
Automatic readability assessment, BiLSTM, attention mechanism, Linguistic-feature fusion, Ordinal classification, RoBERTa, CEFR-SP, OneStopEnglish.
References
- Jinshan Zeng et al., “InterpretARA: Enhancing Hybrid Automatic Readability Assessment with Linguistic Feature Interpreter and Contrastive Learning,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 17, pp. 19497-19505, 2024.
[CrossRef] [Google Scholar] [Publisher Link] - Jinshan Zeng et al., “Enhancing Automatic Readability Assessment with Pre-Training and Soft Labels for Ordinal Regression,” Findings of the Association for Computational Linguistics, pp. 4557-4568, 2022.
[CrossRef] [Google Scholar] [Publisher Link] - Zhenzhen Li, Han Ding, and Shaohong Zhang, “Cross-Corpus Readability Compatibility Assessment for English Texts,” IEEE Access, vol. 11, pp. 101985-101997, 2023.
[CrossRef] [Google Scholar] [Publisher Link] - Fengkai Liu, Tan Jin, and John S. Y. Lee, “Automatic Readability Assessment for Sentences: Neural, Hybrid and Large Language Models,” Language Resources and Evaluation, vol. 59, no. 3, pp. 2265-2296, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Jinshan Zeng et al., “Self-supervised Collaborative Information Bottleneck for Text Readability Assessment,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 24, pp. 25814-25822, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Jinshan Zeng et al., “LearnARA: Automatic Readability Assessment with Deep Linguistic Representation Learner and Contrastive Information Bottleneck,” IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pp. 992-1003, 2026.
[CrossRef] [Google Scholar] [Publisher Link] - Xinying Qiu et al., “Learning Syntactic Dense Embedding with Correlation Graph for Automatic Readability Assessment,” Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, pp. 3013-3025, 2021.
[CrossRef] [Google Scholar] [Publisher Link] - Özkan Aslan, Caner Balım, and Naim Karasekreter, “Advancing Turkish Readability Assessment with Multi-Layer Linguistic Features and Ordinal-Aware Deep Learning,” 2025 International Conference on Intelligent Systems: Theories and Applications (SITA), Rabat, Morocco, pp. 1-6, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Fengkai Liu, and John Lee, “Hybrid Models for Sentence Readability Assessment,” Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications, pp. 448-454, 2023.
[CrossRef] [Google Scholar] [Publisher Link] - Ho Hung Lim, and John Lee, “Improving Readability Assessment with Ordinal Log-Loss,” Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications, Association for Computational Linguistics, Mexico, pp. 343-350, 2024.
[Google Scholar] [Publisher Link] - Ahmet Yavuz Uluslu, and Gerold Schneider, “Exploring Linguistic Features for Turkish Text Readability,” arXiv preprint, pp. 1-10, 2023.
[CrossRef] [Google Scholar] [Publisher Link] - Sepideh Ravanbakhsh, and Mohammad Mahmoodi Varnamkhast, “Persian Text Readability Assessment with Hierarchical Transformer-Based Classification Models,” Scientific Reports, vol. 16, no. 1, pp. 1-11, 2026.
[CrossRef] [Google Scholar] [Publisher Link] - Khalid N. Elmadani, Nizar Habash, and Hanada Taha-Thomure, “A Large and Balanced Corpus for Fine-Grained Arabic Readability Assessment,” Association for Computational Linguistics, pp. 16376-16400, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Justin Lee, and Sowmya Vajjala, “A Neural Pairwise Ranking Model for Readability Assessment,” Findings of the Association for Computational Linguistics, pp. 3802-3813, 2022.
[CrossRef] [Google Scholar] [Publisher Link] - Bruce W. Lee, Yoo Sung Jang, and Jason Lee, “Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features,” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 10669-10686, 2021.
[CrossRef] [Google Scholar] [Publisher Link] - Li Zhang et al., “Automatic Text Readability Assessment for Educational Content Based on Graph Representation Learning,” Scientific Reports, vol. 16, pp. 1-15, 2026.
[CrossRef] [Google Scholar] [Publisher Link] - Suna-Şeyma Uçar et al., “Exploring Automatic Readability Assessment for Science Documents within a Multilingual Educational Context,” International Journal of Artificial Intelligence in Education, vol. 34, no. 4, pp. 1417-1459, 2024.
[CrossRef] [Google Scholar] [Publisher Link] - Wenbiao Li, Wang Ziyang, and Yunfang Wu, “A Unified Neural Network Model for Readability Assessment with Feature Projection and Length-Balanced Loss,” Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 7446-7457, 2022.
[CrossRef] [Google Scholar] [Publisher Link] - Jun Zhao, Readability Assessment of Chinese Linguistic Texts based on Dependent Syntactic Networks, Applied Mathematics and Nonlinear Sciences, vol. 9, no. 1, pp. 1-16, 2024. [Online]. Available: https://www.researchgate.net/publication/378229938_Readability_Assessment_of_Chinese_Linguistic_Texts_Based_on_Dependent_Syntactic_Networks
- Yuchen Wang et al., “Automatically Difficulty Grading Method for English Reading Corpus With Multifeature Embedding Based on a Pretrained Language Model,” IEEE Transactions on Learning Technologies, vol. 17, pp. 474-484, 2024.
[CrossRef] [Google Scholar] [Publisher Link] - Mohamed Amine Ouassil et al., “Enhancing Arabic Text Readability Assessment: A Combined BERT and BiLSTM Approach,” 2024 International Conference on Circuit, Systems and Communication (ICCSC), Fes, Morocco, pp. 1-7, 2024.
[CrossRef] [Google Scholar] [Publisher Link] - Yurui Zheng, Yijun Chen, and Shaohong Zhang, “Hierarchical Ranking Neural Network for Long-Document Readability Assessment,” arXiv preprint, pp. 1-18, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Diego Palma, and Christian Soto, “Combining Large Language Models with Linguistic Features for the Readability Complexity Assessment of Texts,” AHFE International, vol. 199, 2025.
[CrossRef] [Publisher Link] - Catarina Belem et al., “Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics,” Proceedings of the Fourth Workshop on Text Simplification, Accessibility and Readability, Association for Computational Linguistics, Suzhou, China, pp. 47-69, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Zarah Weiss, and Detmar Meurers, “Assessing Sentence Readability for German Language Learners with Broad Linguistic Modelling or Readability Formulas: When do Linguistic Insights Make a Difference?,” Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications, Association for Computational Linguistics, pp. 141-153, 2022.
[CrossRef] [Google Scholar] [Publisher Link] - Sowmya Vajjala, “Trends, Limitations and Open Challenges in Automatic Readability Assessment Research,” Proceedings of the Thirteenth Language Resources and Evaluation Conference, European Language Resources Association, Marseille, France, pp. 5366-5377, 2022.
[Google Scholar] [Publisher Link] - Ziyang Wang et al., “FPT: Feature Prompt Tuning for Few-Shot Readability Assessment,” Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics, Mexico, pp. 280-295, 2024.
[CrossRef] [Google Scholar] [Publisher Link] - Juan Liberato et al., “Strategies for Arabic Readability Modelling,” Proceedings of the Second Arabic Natural Language Processing Conference, Association for Computational Linguistics, Bangkok, Thailand, pp. 55-66, 2024.
[CrossRef] [Google Scholar] [Publisher Link] - Mykola Trokhymovych, Indira Sen, and Martin Gerlach, “An Open Multilingual System for Scoring Readability of Wikipedia,” arXiv preprint, pp. 1-16, 2024.
[CrossRef] [Google Scholar] [Publisher Link] - Nizar Habash et al., “Guidelines for Fine-Grained Sentence-Level Arabic Readability Annotation,” Proceedings of the 19th Linguistic Annotation Workshop, Association for Computational Linguistics, Vienna, Austria, pp. 359-376, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Wenjing Pan et al., “Textual Form Features for Text Readability Assessment,” Natural Language Processing, vol. 31, no. 3, pp. 800-841, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Eugenio Ribeiro, Nuno Mamede, and Jorge Baptista “Automatic Assessment of Text-Complexity Levels in European Portuguese,” Linguamática, vol. 16, no. 2, pp. 115-139, 2024.
[CrossRef] [Publisher Link] - Xiaopeng Zhang, and Xiaofei Lu, “Aligning Linguistic Complexity with the Difficulty of English Texts for L2 Learners Based on CEFR Levels,” Studies in Second Language Acquisition, vol. 47, no. 5, pp. 1407-1434, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Abdul Aziz et al., “Leveraging Contextual Representations with a BiLSTM-based Regressor for Lexical Complexity Prediction,” Natural Language Processing Journal, vol. 5, pp. 1-16, 2023.
[CrossRef] [Google Scholar] [Publisher Link] - D.A. Morozov, A.V. Glazkov, and B.L. Iomdin, “Text Complexity and Linguistic Features: Their Correlation in English and Russian,” Russian Journal of Linguistics, vol. 26, no. 2, pp. 426-448, 2022.
[CrossRef] [Google Scholar] [Publisher Link] - Yuki Arase, Satoru Uchida, and Tomoyuki Kajiwara, “CEFR-based Sentence-Difficulty Annotation and Assessment,” Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, pp. 6206-6219, 2022.
[CrossRef] [Google Scholar] [Publisher Link] - Nour Rabih, “Noor at BAREC Shared Task 2025: A Hybrid Transformer-Feature Architecture for Sentence-level Readability Assessment,” Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks, Association for Computational Linguistics, Suzhou, China, pp. 331-342, 2025.
[Google Scholar] [Publisher Link] - Mutaz Ayesh, “PalNLP at BAREC Shared Task 2025: Predicting Arabic Readability Using Ordinal Regression and K-Fold Ensemble Learning,” Proceedings of The Third Arabic Natural Language Processing Conference, pp. 343-349, 2025.
[Google Scholar] [Publisher Link] - Sandra Aluisio et al., “Readability Assessment for Text Simplification,” Proceedings of the NAACL HLT 2010 Fifth Workshop on Innovative Use of NLP for Building Educational Applications, Association for Computational Linguistics, Los Angeles, California, pp. 1-9, 2010.
[Google Scholar] [Publisher Link] - Rodrigo Wilkens et al., “Exploring Hybrid Approaches to Readability: Experiments on the Complementarity between Linguistic Features and Transformers,” Findings of the Association for Computational Linguistics, pp. 2316-2331, 2024.
[CrossRef] [Google Scholar] [Publisher Link] - Varun Sai Alaparthi et al., “Rating Ease of Readability Using Transformers,” 2022 14th International Conference on Computer and Automation Engineering (ICCAE), Brisbane, Australia, pp. 117-121, 2022.
[CrossRef] [Google Scholar] [Publisher Link] - Susmoy Chakraborty, Mir Tafseer Nayeem, and Wasi Uddin Ahmad, “Simple or Complex? Learning to Predict the Readability of Bengali Texts,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 14, pp. 12621-12629, 2021.
[CrossRef] [Google Scholar] [Publisher Link] - Zijie Zeng, Dragan Gasevic, and Guangliang Chen, “On the Effectiveness of Curriculum Learning in Educational Text Scoring,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 12, pp. 14602-14610, 2023.
[CrossRef] [Google Scholar] [Publisher Link] - Herianah et al., “Automated Assessment of Text Complexity through the Fusion of AutoML and Psycholinguistic Models,” Forum for Linguistic Studies, vol. 7, no. 3, pp. 46-62, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - XiuHua Zhao, “A Hybrid Deep-Learning and Fuzzy-Logic Framework for Feature-Based Evaluation of English-Language Learners,” Scientific Reports, vol. 15, no. 14, pp. 1-40, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Muhammad Zulqarnain, and Muhammad Saqlain, “Text Readability Evaluation in Higher Education using CNNs,” Journal of Industrial Intelligence, vol. 1, no. 3, pp. 184-193, 2023.
[CrossRef] [Google Scholar] [Publisher Link] - Ahmad Jaber Mayahi, and Emman Naser Alshatti, “Assessing English-Language Writing and Readability Skills using a Long Short-Term Memory Model,” 2023 Computer Applications & Technological Solutions (CATS), Mubarak Al-Abdullah, Kuwait, pp. 1-4, 2023.
[CrossRef] [Google Scholar] [Publisher Link] - Parahonco Alexandr, and Parahonco Liudmila, “Feature-Level Decomposition of Text Complexity: Cross-Domain Empirical Evidence,” Computer Science Journal of Moldova, vol. 33, no. 2, pp. 257-280, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - G. Pavani et al., “An Interpretable Dual-Level Feedback Approach for Improving Graded Language Simplification and Readability,” International Journal of Advanced Computer Science and Applications, vol. 16, no. 11, pp. 1-14, 2025.
[CrossRef] [Google Scholar] [Publisher Link] - Yi Zuo, “Automatic Generation of ESL Learning Materials based on CEFR Levels Using Reinforcement-Tuned Large Language Models,” Discover Artificial Intelligence, vol. 6, no. 1, pp. 1-24, 2026.
[CrossRef] [Google Scholar] [Publisher Link]