Machine Learning/ML-এর ভিত্তি/Lesson 03
Features ও Labels
X আর y — model কী দেখে আর কী predict করে। Numerical, binary, ordinal আর nominal feature; one-hot encoding; কোন column feature হতে পারে না (leakage) — আর একটা model-ready feature matrix বানানো।
- সময়
- 24 মিনিট
- Exercise
- 1
- Challenge
- 1
- Quiz
- 3 প্রশ্ন
সহজ ভাষায়
Supervised learning-এ data দুই ভাগে:
- Features (X) — model যা দেখে: পড়ার সময়, উপস্থিতি, tutor…
- Label / target (y) — model যা predict করবে: pass নাকি fail
X (features) y (label)
study_hours attendance tutor parent_education passed
14.2 88 1 2 ──► 1
3.5 62 0 0 ──► 0
9.0 81 1 3 ──► 1
ML convention:
- X — বড় হাতের অক্ষর, কারণ এটা একটা matrix (2D: sample × feature)
- y — ছোট হাতের, কারণ এটা একটা vector (1D: প্রতিটা sample-এর একটা উত্তর)
কেন দরকার?
Model শুধু সংখ্যা বোঝে। "Dhaka", "yes", "secondary" — এগুলো সরাসরি দেওয়া যায় না। কীভাবে সংখ্যায় রূপান্তর করবে, সেই সিদ্ধান্ত model-এর performance-এ বিশাল প্রভাব ফেলে। আর কোন column feature হতে পারবে না — সেই সিদ্ধান্ত ভুল হলে model পুরোপুরি অকেজো।
Feature-এর ধরন
| ধরন | উদাহরণ | কীভাবে সংখ্যা বানাবো |
|---|---|---|
| Numerical | study_hours, attendance | যেমন আছে |
| Binary | tutor (yes/no), gender | 0 / 1 |
| Ordinal (ক্রম আছে) | parent_education: none < primary < secondary < higher | ক্রম অনুযায়ী 0, 1, 2, 3 |
| Nominal (ক্রম নেই) | district | one-hot encoding |
One-hot encoding
District-এর কোনো ক্রম নেই। Dhaka=1, Sylhet=3 লিখলে model ভাববে Sylhet "বড়" — অর্থহীন। সমাধান: প্রতিটা district-এর জন্য আলাদা 0/1 column:
district district_Barishal district_Chattogram district_Dhaka ...
Chattogram 0 1 0
Dhaka 0 0 1
প্রতিটা row-এ ঠিক একটা 1 — "এই student কোন district-এর"। ৮টা district → ৮টা column।
Linear model-এ কখনো কখনো একটা column বাদ দেওয়া হয় (
drop_first=True) — বাকি ৭টা সব 0 মানেই অষ্টমটা। Tree-based model-এ সাধারণত দরকার নেই।
কোন column feature হবে না?
Feature বাছাইয়ের সময় প্রশ্ন: "Prediction করার মুহূর্তে কি এই তথ্য জানা থাকবে?"
| Column | Feature? | কেন |
|---|---|---|
study_hours, attendance, tutor … | ✓ | আগে থেকে জানা |
student_id | ✗ | শুধু পরিচয় — কোনো অর্থ নেই, model ভুল pattern শিখতে পারে |
math, science, english | ✗ | পরীক্ষার পরের তথ্য — leakage |
avg_score | ✗ | এটা দিয়েই passed হিসাব করা — উত্তর দিয়ে উত্তর বলা |
Leakage-এর ফলাফল দেখো:
Leakage-ওয়ালা model "১০০%"! দেখে মনে হয় দারুণ — কিন্তু আসল ব্যবহারে পরীক্ষার আগে avg_score থাকবেই না। এটা ML-এর সবচেয়ে বিপজ্জনক ভুলগুলোর একটা, কারণ ফলাফল ভালো দেখায়।
Exercise
Exercise
X আর y বানাও
students_clean.csv থেকে model-এর জন্য:
X — এই column-গুলো, এই ক্রমে, সব সংখ্যায়:
- study_hours, attendance — যেমন আছে
- tutor, internet — "yes" → 1, "no" → 0
- parent_education — ক্রম অনুযায়ী: "none" → 0, "primary" → 1, "secondary" → 2, "higher" → 3
y — passed column, 0/1 integer হিসেবে।
Quiz
Challenge
Challenge
সম্পূর্ণ feature matrix
make_features(df) function লেখো যেটা যেকোনো student DataFrame থেকে model-ready X বানায়:
study_hours,attendance— যেমন আছেis_female(gender F → 1),is_private(school_type Private → 1),tutor,internet— 0/1parent_education— 0 থেকে 3 (আগের exercise-এর মতো)district— one-hot: প্রতিটা district-এর জন্যdistrict_<নাম>column, 0/1 integer- বাদ দেবে:
student_id(শুধু পরিচয়), আর পরীক্ষার পরের তথ্য —math,science,english,avg_score,passed(leakage!)
সব column সংখ্যা, কোনো missing নেই। মূল df বদলাবে না।
বাস্তবে কোথায় ব্যবহার হয়?
Challenge-এ make_features-কে একটা function বানানোর কারণ — production-এ নতুন student-এর data এলে হুবহু একই রূপান্তর লাগবে। Training-এ এক রকম, prediction-এ আরেক রকম হলে model ভুল উত্তর দেবে, কোনো error ছাড়াই।
একটা সূক্ষ্ম ফাঁদ:
নতুন data-য় মাত্র ২টা district — তাই মাত্র ২টা column! Training-এ ছিল ৮টা। Model-এ দিলে error, বা আরও খারাপ, column ভুল জায়গায়। Feature Engineering lesson-এ scikit-learn-এর OneHotEncoder আর Pipeline দিয়ে এর সঠিক সমাধান দেখবো — যেটা training-এর category মনে রাখে।
Interview প্রশ্ন
- Beginner: Feature আর label কী? X বড় হাতের আর y ছোট হাতের কেন?
- Intermediate: One-hot আর ordinal encoding-এর পার্থক্য কী? কখন কোনটা?
- Advanced: Target leakage কী? একটা real উদাহরণ দাও যেখানে leakage ধরা কঠিন। (যেমন hospital data-য় "chemotherapy দেওয়া হয়েছে কিনা" দিয়ে cancer predict করা।)
এরপর কী?
X আর y তৈরি। এখন প্রথম আসল model train করার পালা — আর সবচেয়ে জরুরি প্রশ্নের উত্তর: "model কি নতুন student-দের ক্ষেত্রেও ঠিক কাজ করবে?" পরের lesson: Train/Test split ও first model।