Machine Learning Case Study

Product PricePrediction

An end-to-end machine learning project that explores e-commerce pricing data, engineers predictive features, and benchmarks regression algorithms to estimate product prices from historical marketplace observations.

306K+ observations26 featuresRegression benchmarking

Price Intelligence

Exploratory data analysis

EDA

Original records

303,982

Cleaned records

281,147

Data preparation

01

Exploration

Complete
02

Cleaning

Complete
03

Feature engineering

Complete
04

Model benchmarking

Evaluated

Historical marketplace data · January–March 2025

01 / Overview

Understanding the pricing problem

Accurate price estimation requires more than fitting a regression model. The project starts by understanding the marketplace data, validating its quality, and identifying which product attributes carry meaningful pricing information.

Dataset size

306,226

Observations in the original analytical dataset.

Features

26

Product, shop, pricing, promotion, and temporal attributes.

Time coverage

81 days

January 1 to March 22, 2025.

After cleaning

281,147

Records retained after data quality checks.

The objective

Develop a reproducible workflow for estimating product prices from historical structured data, then evaluate how different regression algorithms respond to numerical and categorical features.

The data challenges

Missing collection dates, duplicate records, inconsistent promotion information, invalid prices, outliers, and relationships between product identifiers all need to be investigated before modeling.

02 / Exploratory Data Analysis

Understanding the distribution of prices

The distribution of the target variable influences how a regression model learns. These visualizations compare raw product prices with their log1p-transformed values before modeling.

01

Original Price Distribution

Raw product prices have a strongly right-skewed distribution. A small number of expensive products can stretch the scale and make the majority of observations harder to analyze.

Histogram showing the original product price distribution
02

Log1p Price Distribution

Applying log1p compresses extreme values while preserving the ordering of prices. This produces a more manageable target distribution for regression experiments.

Histogram showing product prices after log1p transformation

Why use log1p?

The transformation log1p(price) computes the natural logarithm of one plus the price. It reduces the influence of very large values on the target scale, making the distribution easier to inspect and potentially improving regression performance. Predictions in transformed space must be converted back with expm1 when prices in the original scale are required.

03 / Feature Analysis

Finding meaningful pricing signals

Beyond the target distribution, the project examines relationships between product attributes, historical prices, promotions, and marketplace identifiers to understand which variables may help explain price variation.

1

Historical price attributes

item_price_min, item_price_max, and priceBeforeDiscount are identified as strong numerical predictors.

2

Product identifiers

itemId and modelId capture product-level information that broader categories may not represent.

3

Promotion integrity

Discount calculations and promotional attributes are checked for consistency and anomalies.

4

Temporal behavior

Observation dates, weekend effects, and price stability help characterize changes over time.

Feature relationship graph

Visual analysis of relationships in the dataset

Graph showing relationships between product features and price

04 / Methodology

From raw data to model evaluation

The workflow is organized into notebook-based experiments, separating data exploration and preparation from regression baselines and categorical feature benchmarking.

Step 01

Exploratory Data Analysis

Investigate data quality, temporal coverage, price distributions, product relationships, promotions, and potential anomalies.

Step 02

Data Cleaning

Remove duplicate records, handle missing values, standardize categorical values, validate discounts, and filter invalid prices.

Step 03

Feature Engineering

Prepare temporal features, encode categorical variables, scale numerical features, and transform the target with log1p.

Step 04

Baseline Regression

Benchmark linear and tree-based regression algorithms using MAE, RMSE, MAPE, R², and training time.

Step 05

Categorical Benchmarking

Compare XGBoost, LightGBM, and CatBoost using individual product and marketplace categorical features.

Before cleaning

303,982

Records entering the cleaning stage.

After cleaning

281,147

Records remaining after quality assurance.

05 / Experimental Results

Benchmarking regression models

The project evaluates classical regression baselines and compares three gradient boosting frameworks. The following tables show the documented categorical feature experiments, with the remaining numerical features kept in the experiment.

Best reported MAE

1.92

XGBoost using modelId.

Best reported RMSE

21.05

XGBoost using modelId.

Best reported MAPE

0.41%

XGBoost using modelId.

Reported R²

0.9995

XGBoost using modelId.

An important modeling insight

Across the documented categorical feature experiments, modelId provides the strongest reported results. This suggests that product-level identifiers contain substantial pricing information. However, the exceptionally high R² should be interpreted alongside the data split strategy and potential identifier leakage before assuming the same performance will generalize to unseen products.

XGBoost

Categorical feature comparison with remaining numerical features.

FeatureMAERMSEMAPER²
modelIdTop result1.9221.050.41%0.9995
itemId25.9899.777.72%0.9885
cat_id28.91118.317.82%0.9838
brand28.69126.117.84%0.9816
promotionId31.26146.717.93%0.9752

LightGBM

Gradient boosting benchmark across individual categorical features.

FeatureMAERMSEMAPER²
modelIdTop result7.0453.521.30%0.9967
itemId26.49102.227.56%0.9879
brand27.40101.038.00%0.9882
cat_id27.68100.878.02%0.9883
promotionId27.80100.678.11%0.9883

CatBoost

Alternative boosting model evaluated using the same feature groups.

FeatureMAERMSEMAPER²
modelIdTop result25.8391.946.65%0.9902
promotionId30.49108.218.47%0.9865
itemId30.55108.188.47%0.9865
brand30.60107.768.51%0.9866
cat_id30.62108.328.46%0.9865

How to interpret the metrics

MAE

Mean Absolute Error measures the average absolute difference between predictions and actual values.

RMSE

Root Mean Squared Error penalizes larger prediction errors more heavily.

MAPE

Mean Absolute Percentage Error expresses errors relative to actual values, with limitations around zero-valued targets.

R²

The coefficient of determination measures how much target variation is explained by the model on the evaluated data.

06 / Key Takeaways

What the experiments reveal

The findings emphasize data quality, target transformation, and the importance of selecting meaningful features before increasing model complexity.

01

Feature quality matters

Historical price attributes and product-level identifiers provide strong predictive signals in the evaluated dataset.

02

Understand the target distribution

Log1p transformation reduces skewness and can make the target easier to model.

03

Benchmark multiple algorithms

Linear models and tree-based ensembles provide useful baselines for understanding nonlinear pricing patterns.

04

Validate generalization

Time-aware validation and product-level holdouts help assess whether a model predicts future prices or genuinely unseen products.

07 / Technology

Tools behind the project

A Python-based analytical workflow built around data manipulation, visualization, and classical machine learning.

PythonPandasNumPyMatplotlibScikit-learnXGBoostLightGBMCatBoostJupyter Notebook

Explore the implementation

From exploratory analysis to machine learning benchmarks.

Explore the notebooks, preprocessing steps, and model experiments in the project repository.

Source Code