🐍 Advanced Python Programming — Networkingrequests · HTTP · CSV · pandas pipeline
SINGLE HTML · COLAB GUIDE
ADVANCED PYTHON · NETWORKING

Exercise 3 · HTTP CSV Data Processing

This page is a standalone code guide. Copy each Python block to Google Colab and run it there.

Remote Data Pipeline

Download CSV data through HTTP, validate the response, then convert the returned text into a pandas DataFrame.

URL→HTTP GET→CSV Text→DataFrame
1. Send request with requests.get()
2. Validate with raise_for_status()
3. Read response.text
4. Parse CSV using pandas
Run this block first in Colab. It creates df, which Step 2 uses.

Python Code · Download + Parse

import requests
import pandas as pd
from io import StringIO

url = "https://raw.githubusercontent.com/cs109/2014_data/master/countries.csv"

response = requests.get(url, timeout=10)
response.raise_for_status()

print("HTTP:", response.status_code)
print("Content-Type:", response.headers.get("content-type"))

df = pd.read_csv(StringIO(response.text))

print(df.head())
print("Rows:", len(df))

Python Code · Process Data

# Continue from df created in Step 1

print("\nMissing values")
print(df.isna().sum())

print("\nCount by Region")
summary = (
    df.groupby("Region")
      .size()
      .sort_values(ascending=False)
)
print(summary)

# Filter one region
europe = df[df["Region"] == "EUROPE"]

print("\nEurope rows:", len(europe))
print(europe.head())

Mission

  • Count downloaded records
  • Inspect missing values
  • Group records by Region
  • Filter only EUROPE
  • Separate HTTP/network work from data-processing work
Colab order: Run Tab 1 code first → then run Tab 2 code in another Colab cell.
Challenge
Replace the URL with another public CSV dataset. Keep the HTTP pipeline but modify the processing logic to match the new columns.