{ "nbformat": 4, "nbformat_minor": 0, "metadata": { "colab": { "name": "Determinants of Earnings (start).ipynb", "provenance": [], "toc_visible": true }, "kernelspec": { "name": "python3", "display_name": "Python 3" } }, "cells": [ { "cell_type": "markdown", "metadata": { "id": "BHg0HZz-intQ" }, "source": [ "# Introduction" ] }, { "cell_type": "markdown", "metadata": { "id": "V2RQkgAbiqJv" }, "source": [ "The National Longitudinal Survey of Youth 1997-2011 dataset is one of the most important databases available to social scientists working with US data. \n", "\n", "It allows scientists to look at the determinants of earnings as well as educational attainment and has incredible relevance for government policy. It can also shed light on politically sensitive issues like how different educational attainment and salaries are for people of different ethnicity, sex, and other factors. When we have a better understanding how these variables affect education and earnings we can also formulate more suitable government policies. \n", "\n", "
\n" ] }, { "cell_type": "markdown", "metadata": { "id": "YjCPWWUSirY_" }, "source": [ "### Upgrade Plotly" ] }, { "cell_type": "code", "metadata": { "id": "v74l3QCGirIX" }, "source": [ "%pip install --upgrade plotly" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "br_QkHBMjC1Q" }, "source": [ "### Import Statements\n" ] }, { "cell_type": "code", "metadata": { "id": "gSKZx-kwie_u" }, "source": [ "import pandas as pd\n", "import numpy as np\n", "\n", "import seaborn as sns\n", "import plotly.express as px\n", "import matplotlib.pyplot as plt\n", "\n", "from sklearn.linear_model import LinearRegression\n", "from sklearn.model_selection import train_test_split" ], "execution_count": 3, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "9pgsrth_izCn" }, "source": [ "## Notebook Presentation" ] }, { "cell_type": "code", "metadata": { "id": "Cgwu-WbBizqY" }, "source": [ "pd.options.display.float_format = '{:,.2f}'.format" ], "execution_count": 4, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "E5bhysOOjLRr" }, "source": [ "# Load the Data\n", "\n" ] }, { "cell_type": "code", "metadata": { "id": "6VngeTQwjM-X" }, "source": [ "df_data = pd.read_csv('NLSY97_subset.csv')" ], "execution_count": 6, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "_ZjIBJ5jjrj0" }, "source": [ "### Understand the Dataset\n", "\n", "Have a look at the file entitled `NLSY97_Variable_Names_and_Descriptions.csv`. \n", "\n", "---------------------------\n", "\n", " :Key Variables: \n", " 1. S Years of schooling (highest grade completed as of 2011)\n", " 2. EXP Total out-of-school work experience (years) as of the 2011 interview.\n", " 3. EARNINGS Current hourly earnings in $ reported at the 2011 interview" ] }, { "cell_type": "markdown", "metadata": { "id": "8MkSxkjVnIfW" }, "source": [ "# Preliminary Data Exploration 🔎\n", "\n", "**Challenge**\n", "\n", "* What is the shape of `df_data`? \n", "* How many rows and columns does it have?\n", "* What are the column names?\n", "* Are there any NaN values or duplicates?" ] }, { "cell_type": "code", "metadata": { "id": "V_cQguBbjwZv" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "4itxqgP6nQj3" }, "source": [ "## Data Cleaning - Check for Missing Values and Duplicates\n", "\n", "Find and remove any duplicate rows." ] }, { "cell_type": "code", "metadata": { "id": "J3DHEFXWnS2N" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "gGmBBPxZnVKC" }, "source": [ "## Descriptive Statistics" ] }, { "cell_type": "code", "metadata": { "id": "I5VP2BMVnVrt" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "ZO-86NXbnWSH" }, "source": [ "## Visualise the Features" ] }, { "cell_type": "code", "metadata": { "id": "hFZJjbsKncPM" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "9i4zHYG4nhDL" }, "source": [ "# Split Training & Test Dataset\n", "\n", "We *can't* use all the entries in our dataset to train our model. Keep 20% of the data for later as a testing dataset (out-of-sample data). " ] }, { "cell_type": "code", "metadata": { "id": "M_OfRSyunkA1" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "EM99NOH0noFS" }, "source": [ "# Simple Linear Regression\n", "\n", "Only use the years of schooling to predict earnings. Use sklearn to run the regression on the training dataset. How high is the r-squared for the regression on the training data? " ] }, { "cell_type": "code", "metadata": { "id": "J_MViuoNnvHf" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "s2TeWKs7oJSa" }, "source": [ "### Evaluate the Coefficients of the Model\n", "\n", "Here we do a sense check on our regression coefficients. The first thing to look for is if the coefficients have the expected sign (positive or negative). \n", "\n", "Interpret the regression. How many extra dollars can one expect to earn for an additional year of schooling?" ] }, { "cell_type": "code", "metadata": { "id": "QmhzZAmAoW4t" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "code", "metadata": { "id": "e9hpdAt3oWnq" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "WIyMPXXYobx8" }, "source": [ "### Analyse the Estimated Values & Regression Residuals\n", "\n", "How good our regression is also depends on the residuals - the difference between the model's predictions ( 𝑦̂ 𝑖 ) and the true values ( 𝑦𝑖 ) inside y_train. Do you see any patterns in the distribution of the residuals?" ] }, { "cell_type": "code", "metadata": { "id": "khkgscweosP_" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "code", "metadata": { "id": "m_diDXSXotm6" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "code", "metadata": { "id": "6DfAEUWNosHd" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "zNBuJ1iBnvpl" }, "source": [ "# Multivariable Regression\n", "\n", "Now use both years of schooling and the years work experience to predict earnings. How high is the r-squared for the regression on the training data? " ] }, { "cell_type": "code", "metadata": { "id": "Ihq-C4looCSM" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "code", "metadata": { "id": "dRhB7Iwboyfq" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "nHDtunM0oyuk" }, "source": [ "### Evaluate the Coefficients of the Model" ] }, { "cell_type": "code", "metadata": { "id": "5vasqInIoydB" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "Yv8q90IYou2Q" }, "source": [ "### Analyse the Estimated Values & Regression Residuals" ] }, { "cell_type": "code", "metadata": { "id": "8NmXnsxfowkI" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "code", "metadata": { "id": "0ZZ1e0spo5o1" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "YWNeoqPLpjVb" }, "source": [ "# Use Your Model to Make a Prediction\n", "\n", "How much can someone with a bachelors degree (12 + 4) years of schooling and 5 years work experience expect to earn in 2011?" ] }, { "cell_type": "code", "metadata": { "id": "Mof-14lCpv60" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "code", "metadata": { "id": "3htX8_SBpvyb" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "markdown", "metadata": { "id": "TIYI-eQepDSQ" }, "source": [ "# Experiment and Investigate Further\n", "\n", "Which other features could you consider adding to further improve the regression to better predict earnings? " ] }, { "cell_type": "code", "metadata": { "id": "sd07-pKopJgo" }, "source": [ "" ], "execution_count": null, "outputs": [] }, { "cell_type": "code", "metadata": { "id": "Fohe2-Rdp1MO" }, "source": [ "" ], "execution_count": null, "outputs": [] } ] }