Retrieval-oriented pre-training outperforms language specialization: Evidence from Bengali university FAQ retrieval
| bracu.degree.level | Postgraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Sadeque, Farig Yousuf | |
| dc.contributor.author | Arman, Mithila | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-08-30T08:49:10Z | |
| dc.date.available | 2026-08-30T08:49:10Z | |
| dc.date.copyright | 2026 | |
| dc.date.issued | 2026-05 | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2026. | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 59-63). | |
| dc.description.abstract | Approaching Bengali institutional FAQ answering as a dense bi-encoder retrieval task, this study introduces University FAQs, a real-world benchmark. The dataset, extracted from 10,264 raw records, is cleaned up to obtain data on 7,233 question– answer pairs distributed among 148 topic labels that allow us the systematic evaluation of a total of 18 sentence encoders including Bengali specific sentence encoders, multilingual SBERT variants and E5 retrievers the BGE-M3 as well as GTE considered in this work along with several MLM based baselines under a unified fine-tuning framework with Multiple Negatives Ranking Loss. This study moves beyond traditional retrieval metrics by also evaluating ranking behaviour, semanticspace quality, calibration, efficiency and levels of resistance to character-level noise with a view to defining explicit realistic deployment requirements of educational support systems. On a broad level, the results demonstrate retrieval-oriented multilingual contrastive encoders outperform Bengali specific and MLM-only baselines signal, clearly and consistently. Out of all models, multilingual-e5-large performs the best overall (Acc@1 = 0.923, MRR@10 = 0.949), with BGE-M3 and multilinguale5- large-instruct in close second and third place respectively. This study analysis also demonstrates that contrastive pre-training leads to a mean Acc@1 gain of 0.175 over MLM-only models, whilst multilingual-e5-small provides the best quality– efficiency trade-off in latency-sensitive settings. Moreover, a confidence-based abstention strategy yields 96.4% accuracy while only answering 68.2% of queries. More generally, this work lays a solid empirical foundation for Bengali FAQ retrieval and shows that multilingual retrieval-oriented pre-training may be more beneficial than | |
| dc.description.degree | Master of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Mithila Arman | |
| dc.format.extent | 86 pages | |
| dc.identifier.other | ID 23266024 | |
| dc.identifier.uri | https://hdl.handle.net/10361/29605 | |
| dc.language.iso | en_US | |
| dc.publisher | BRAC University | |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Bengali FAQ retrieval | |
| dc.subject | Bi-eencoder retrieval | |
| dc.subject | Contrastive learning | |
| dc.subject | Educational question answering | |
| dc.subject | NLP | |
| dc.subject | Bengali language technology | |
| dc.subject.lcsh | Contrastive linguistics. | |
| dc.subject.lcsh | Bengali language--Data processing. | |
| dc.subject.lcsh | Computational linguistics. | |
| dc.subject.lcsh | Natural language generation (Computer science). | |
| dc.title | Retrieval-oriented pre-training outperforms language specialization: Evidence from Bengali university FAQ retrieval | |
| dc.type | Thesis |