REST API Data Retrieval
Fetch and analyze data from public REST APIs using Python and pandas
Objectives
By the end of this practical work, you will be able to:
- Understand REST API concepts (endpoints, HTTP methods, JSON)
- Make HTTP GET requests using Python's
requestslibrary - Parse JSON responses and extract relevant data
- Handle pagination in API responses
- Transform API data into pandas DataFrames for analysis
- Integrate REST API data with Orange Data Mining
Prerequisites
- Python 3.8+ installed
- Basic understanding of JSON format
- Orange Data Mining (optional, for visualization)
Install required packages:
pip install requests pandas
API Selection
We'll work with the REST Countries API, a free public API:
- REST Countries API - Country information
Note: Public APIs don't require API keys, making them perfect for learning!
Instructions
Step 1: Make Your First API Request
Start by fetching data for a single country:
import requests
# Fetch data for France
url = "https://restcountries.com/v3.1/name/france"
response = requests.get(url)
# Check response status
print(f"Status Code: {response.status_code}")
print(f"Content Type: {response.headers.get('content-type')}")
# Parse JSON response
if response.status_code == 200:
data = response.json()
print(f"Number of results: {len(data)}")
else:
print(f"Error: {response.status_code}")
Step 2: Explore the JSON Structure
Understand the nested JSON structure:
import json
# Pretty print the first result
country = data[0]
print(json.dumps(country, indent=2)[:1000])
# Access nested data
print(f"\nCountry: {country['name']['common']}")
print(f"Capital: {country['capital'][0]}")
print(f"Population: {country['population']:,}")
print(f"Area: {country['area']:,} km²")
print(f"Region: {country['region']}")
Step 3: Fetch All Countries
Get data for all countries at once:
# Fetch all countries
url = "https://restcountries.com/v3.1/all"
response = requests.get(url)
all_countries = response.json()
print(f"Total countries: {len(all_countries)}")
# Preview first 5 country names
for country in all_countries[:5]:
name = country['name']['common']
pop = country.get('population', 0)
print(f" - {name}: {pop:,} people")
Step 4: Extract Data to Dictionary
Create a function to extract key fields:
def extract_country_data(country):
"""Extract relevant fields from a country JSON object."""
capitals = country.get('capital', ['N/A'])
capital = capitals[0] if capitals else 'N/A'
languages = country.get('languages', {})
lang_list = list(languages.values()) if languages else []
return {
'name': country['name']['common'],
'capital': capital,
'region': country.get('region', 'Unknown'),
'subregion': country.get('subregion', 'Unknown'),
'population': country.get('population', 0),
'area': country.get('area', 0),
'languages': ', '.join(lang_list),
'landlocked': country.get('landlocked', False)
}
# Test with one country
france_data = extract_country_data(all_countries[0])
for key, value in france_data.items():
print(f"{key}: {value}")
Step 5: Convert to DataFrame
Process all countries into a pandas DataFrame:
import pandas as pd
# Extract data for all countries
countries_data = [extract_country_data(c) for c in all_countries]
# Create DataFrame
df = pd.DataFrame(countries_data)
# Display basic info
print(f"DataFrame shape: {df.shape}")
print(f"\nColumns: {list(df.columns)}")
# Preview the data
print(f"\nFirst 10 countries:")
print(df[['name', 'capital', 'population', 'region']].head(10))
Step 6: Basic Data Analysis
Analyze the country data:
# Population statistics
print("=== Population Statistics ===")
print(f"Total world population: {df['population'].sum():,}")
print(f"Average population: {df['population'].mean():,.0f}")
# Top 10 most populous countries
print("\n=== Top 10 Most Populous Countries ===")
top_10 = df.nlargest(10, 'population')[['name', 'population', 'region']]
print(top_10.to_string(index=False))
# Countries by region
print("\n=== Countries by Region ===")
region_counts = df['region'].value_counts()
print(region_counts)
Step 7: Save to CSV
Export the data for later use:
# Save all countries
df.to_csv('countries_data.csv', index=False)
print(f"Saved {len(df)} countries to countries_data.csv")
Step 8: Integrate with Orange
Create an Orange Data Table from the API data:
from Orange.data import Table, Domain, StringVariable, ContinuousVariable, DiscreteVariable
# Define the domain
domain = Domain(
[ContinuousVariable("population"),
ContinuousVariable("area")],
[DiscreteVariable("region", values=list(df['region'].unique()))],
[StringVariable("name"),
StringVariable("capital"),
StringVariable("languages")]
)
# Prepare data as list of lists
data_list = []
for _, row in df.iterrows():
data_list.append([
row['population'],
row['area'],
row['region'],
row['name'],
row['capital'],
row['languages']
])
# Create Orange table
out_data = Table.from_list(domain, data_list)
print(f"Created Orange table with {len(out_data)} rows")
Expected Output
After completing this practical work, you should have:
- A working Python script that fetches data from REST Countries API
- A CSV file with data for 250+ countries
- Basic statistical analysis of world population and geography
- An Orange Data Table ready for visual analysis
Deliverables
- Python Script: Complete API data retrieval script (.py file)
- CSV Export: countries_data.csv with all extracted data
- Report: Answer these questions:
- How many countries are UN members?
- Which region has the highest total population?
- What is the largest landlocked country by area?
Bonus Challenges
- Challenge 1: Use the Open-Meteo API to fetch weather data for capitals
- Challenge 2: Calculate population density and find most/least dense countries
- Challenge 3: Create a world map visualization
- Challenge 4: Combine with World Bank API for GDP data
API Endpoints Reference
| Endpoint | Description |
|---|---|
/v3.1/all |
Get all countries |
/v3.1/name/{name} |
Search by country name |
/v3.1/region/{region} |
Filter by region |
/v3.1/alpha/{code} |
Get by country code |