Machine Learning / Machine Learning — Regression
Car Price Prediction
A machine learning model that predicts used-car selling prices from vehicle and ownership-related features.
- Python
- Pandas
- NumPy
- Scikit-learn
- XGBoost
- Matplotlib
- Seaborn
Test R²
5-fold CV R²
Unique records
Problem
Used-car pricing depends on a mix of vehicle, ownership, and usage features that need consistent preprocessing before regression.
Approach
The project removes duplicate records, performs EDA and feature engineering, builds a preprocessing pipeline with One-Hot Encoding, applies a log transformation to the target, and trains an XGBoost Regressor.
Architecture
Capabilities
- Car brand / model
- Car age
- Kilometers driven
- Fuel type
- Seller type
- Transmission
- Previous owners
- Kilometers driven per year
Dataset
The dataset contains 4,340 original records and 3,577 unique records after duplicate removal.
Work performed
The project includes EDA, data manipulation and preprocessing, a car age feature, car model / brand features, and kilometers driven per year.
Modeling
Categorical features use One-Hot Encoding in a preprocessing pipeline. The target is log-transformed before training an XGBoost Regressor.
Evaluation
The model was evaluated with R² and MAE, with 5-fold cross-validation used to assess generalization.
Model artifact
The serialized model artifact is car_price_model.pkl.
Outcome
Test R²: 0.793 · 5-fold Cross-Validation R²: 0.777