Corpus Research Center, Minhaj University Lahore
|Research Intern
Lahore, Punjab, Pakistan
Summary
Led contributions to a sentiment analysis project by building a mini Hausa corpus and ensuring data quality through meticulous annotation and collaborative efforts.
Highlights
Contributed significantly to the "Building a Mini Hausa Corpus of Social Media Texts for Sentiment Analysis" project, enhancing linguistic data resources for AI model development.
Collected and meticulously compiled diverse Hausa Facebook posts and comments using Chrome-based extraction tools and manual techniques, ensuring comprehensive data gathering and preprocessing.
Annotated over 1,000 texts for sentiment polarity and linguistic features, directly supporting supervised machine learning tasks and improving model accuracy by an estimated 15%.
Collaborated effectively with multilingual researchers on corpus design, annotation consistency, and quality assurance, upholding high data integrity standards.
Supported documentation and data management, aligning with research ethics and open-source standards to ensure project transparency and reproducibility.