Comparative Evaluation of Naive Bayes and Support Vector Machine Classifiers for Detecting Scam Text Messages in Filipino and English
Scam texts mix Filipino, English and Taglish, which can confuse classifiers trained only on English. Comparing two well-understood models on a local dataset is a contained, reproducible thesis.
- Research type
- Experimental
- Suggested method
- Experimental comparison with k-fold cross-validation on a labeled SMS dataset, ≈3,000 messages
- Feasibility
- High feasibility. Runs on a laptop; collecting and labeling enough real scam messages is the slow part.
- Ethics
- Remove phone numbers and personal details from collected messages before labeling.