Retrieval-oriented pre-training outperforms language specialization: Evidence from Bengali university FAQ retrieval

bracu.degree.levelPostgraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorSadeque, Farig Yousuf
dc.contributor.authorArman, Mithila
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-08-30T08:49:10Z
dc.date.available2026-08-30T08:49:10Z
dc.date.copyright2026
dc.date.issued2026-05
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2026.
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 59-63).
dc.description.abstractApproaching Bengali institutional FAQ answering as a dense bi-encoder retrieval task, this study introduces University FAQs, a real-world benchmark. The dataset, extracted from 10,264 raw records, is cleaned up to obtain data on 7,233 question– answer pairs distributed among 148 topic labels that allow us the systematic evaluation of a total of 18 sentence encoders including Bengali specific sentence encoders, multilingual SBERT variants and E5 retrievers the BGE-M3 as well as GTE considered in this work along with several MLM based baselines under a unified fine-tuning framework with Multiple Negatives Ranking Loss. This study moves beyond traditional retrieval metrics by also evaluating ranking behaviour, semanticspace quality, calibration, efficiency and levels of resistance to character-level noise with a view to defining explicit realistic deployment requirements of educational support systems. On a broad level, the results demonstrate retrieval-oriented multilingual contrastive encoders outperform Bengali specific and MLM-only baselines signal, clearly and consistently. Out of all models, multilingual-e5-large performs the best overall (Acc@1 = 0.923, MRR@10 = 0.949), with BGE-M3 and multilinguale5- large-instruct in close second and third place respectively. This study analysis also demonstrates that contrastive pre-training leads to a mean Acc@1 gain of 0.175 over MLM-only models, whilst multilingual-e5-small provides the best quality– efficiency trade-off in latency-sensitive settings. Moreover, a confidence-based abstention strategy yields 96.4% accuracy while only answering 68.2% of queries. More generally, this work lays a solid empirical foundation for Bengali FAQ retrieval and shows that multilingual retrieval-oriented pre-training may be more beneficial than
dc.description.degreeMaster of Science in Computer Science and Engineering
dc.description.statementofresponsibilityMithila Arman
dc.format.extent86 pages
dc.identifier.otherID 23266024
dc.identifier.urihttps://hdl.handle.net/10361/29605
dc.language.isoen_US
dc.publisherBRAC University
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectBengali FAQ retrieval
dc.subjectBi-eencoder retrieval
dc.subjectContrastive learning
dc.subjectEducational question answering
dc.subjectNLP
dc.subjectBengali language technology
dc.subject.lcshContrastive linguistics.
dc.subject.lcshBengali language--Data processing.
dc.subject.lcshComputational linguistics.
dc.subject.lcshNatural language generation (Computer science).
dc.titleRetrieval-oriented pre-training outperforms language specialization: Evidence from Bengali university FAQ retrieval
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
23266024_CSE.pdf
Size:
3.43 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: