IT 544 Data Visualization · Walsh College · Winter 2025
Her plan for analyzing U.S. pollution data (trends, seasonality, hot days and picking the right chart for each audience) rebuilt as a live workbench on a public daily ozone dataset.
One switch controls missing data for everything below. Watch which results move and which do not.
Her final project planned visual analysis of the Kaggle U.S. Pollution dataset (about 1.4 million rows of PM2.5, NO2, O3, CO and SO2 readings by city, 2000 to 2016). That file is not bundled here, so this page re-creates the same workflow on R's public airquality dataset: 153 days of ozone and weather readings for New York, May to September 1973. The methods carry over; the numbers describe New York in 1973, not her Kaggle data.
… days above the threshold
Illustrative only: 70 ppb is the level of today's EPA 8-hour ozone standard, but these readings are 1 pm to 3 pm averages, so the count is not a regulatory exceedance count.
Computed from the observed readings (missing days dropped) unless the chart says otherwise.
Her document also outlines Random Forest and ARIMA models on the Kaggle file; those results are not reproduced here because that data is not available to this page.
Columns: Ozone = mean ozone in parts per billion, 1 pm to 3 pm, Roosevelt Island. Solar.R = solar radiation in Langleys, 8 am to noon, Central Park. Wind = average wind speed in mph at 7 am and 10 am, LaGuardia Airport. Temp = daily maximum temperature in °F, LaGuardia Airport. Source: New York State Department of Conservation (ozone) and National Weather Service (weather), via R's datasets package and the Rdatasets archive.