All work

Machine Learning / Machine Learning — Regression

Car Price Prediction

A machine learning model that predicts used-car selling prices from vehicle and ownership-related features.

  • Python
  • Pandas
  • NumPy
  • Scikit-learn
  • XGBoost
  • Matplotlib
  • Seaborn
Model snapshot regression
0.793Test R²
0.7775-fold CV R²
3,577Unique records
XGBoost Regressor
0.793

Test R²

0.777

5-fold CV R²

3,577

Unique records

Problem

Used-car pricing depends on a mix of vehicle, ownership, and usage features that need consistent preprocessing before regression.

Approach

The project removes duplicate records, performs EDA and feature engineering, builds a preprocessing pipeline with One-Hot Encoding, applies a log transformation to the target, and trains an XGBoost Regressor.

Architecture

014,340 original records
023,577 unique records after duplicate removal
03Feature engineering
04One-Hot Encoding pipeline
05Log-transformed target
06XGBoost Regressor
07R² and MAE evaluation
085-fold cross-validation

Capabilities

  • Car brand / model
  • Car age
  • Kilometers driven
  • Fuel type
  • Seller type
  • Transmission
  • Previous owners
  • Kilometers driven per year

Dataset

The dataset contains 4,340 original records and 3,577 unique records after duplicate removal.

Work performed

The project includes EDA, data manipulation and preprocessing, a car age feature, car model / brand features, and kilometers driven per year.

Modeling

Categorical features use One-Hot Encoding in a preprocessing pipeline. The target is log-transformed before training an XGBoost Regressor.

Evaluation

The model was evaluated with R² and MAE, with 5-fold cross-validation used to assess generalization.

Model artifact

The serialized model artifact is car_price_model.pkl.

Outcome

Test R²: 0.793 · 5-fold Cross-Validation R²: 0.777