Analytics & Predictive Modeling Track

Data Science Career Roadmap

Master SQL, Python, Statistical Hypothesis Testing, A/B Testing, Exploratory Data Analysis (EDA), Predictive ML, Business Storytelling, and Executive Dashboards.

Step-by-Step Learning Stages

1 Weeks 1 - 4

SQL Mastery & Data Wrangling in Python

Master SQL aggregations, window functions (`RANK`, `DENSE_RANK`, `LAG`), CTEs, joins, and Python Pandas/NumPy dataframe manipulation.

Advanced SQL Pandas DataFrames Data Cleaning
2 Weeks 5 - 8

Statistics, Probability & A/B Testing

Understand Central Limit Theorem, p-values, t-tests, chi-square tests, confidence intervals, sample size estimation, and experimental A/B testing design.

Hypothesis Testing A/B Testing p-Values & t-Tests
3 Weeks 9 - 12

Exploratory Data Analysis & Data Visualization

Build interactive charts and business dashboards with Seaborn, Plotly, Tableau, and Power BI. Tell compelling data stories to executive stakeholders.

Seaborn / Plotly Tableau / PowerBI Data Storytelling
4 Weeks 13 - 18

Machine Learning & Time Series Forecasting

Apply predictive modeling (Regression, Logistic Classification, Random Forests, XGBoost), customer clustering (K-Means, RFM), and time-series forecasting (ARIMA, Prophet).

Scikit-Learn Prophet / ARIMA RFM Analysis

Recommended Portfolio Projects

3 practical projects to prove your analytical & predictive data science skills.

BEGINNER

Executive Sales & Revenue Insights Dashboard

Clean complex retail transaction data with SQL, extract YoY sales growth, customer acquisition costs, and build an interactive Tableau/PowerBI dashboard.

Stack: SQL, Tableau / PowerBI, Excel, PostgreSQL
INTERMEDIATE

E-Commerce A/B Test & Conversion Rate Study

Analyze clickstream experiment data, verify sample ratio mismatch (SRM), compute two-tailed t-tests and p-values, and present business recommendations.

Stack: Python, Statsmodels, SciPy, Plotly, Pandas
ADVANCED

Customer Segmentation & Lifetime Value Engine

Combine Recency-Frequency-Monetary (RFM) scoring with K-Means clustering to identify high-value customer segments and predict 12-month Customer Lifetime Value (CLV).

Stack: Scikit-Learn, K-Means, XGBoost, Seaborn