Use only pure Python and the built-in
csvmodule.
Practice dataset loading, filtering, aggregation, and complexity analysis.
Write your code in:
session1/solutions/exercise-01-homework.pyUse the same solutions/ folder created in Part 1.
Use: movies.csv from Birkbeck/movies
- Load the dataset with
csv.reader. - Print:
- the number of data rows (excluding header)
- the number of columns
- Print the first 3 rows (including header).
- Find and print the first movie where the
genrescolumn containsAction. - Compute and print the average of
rating_imdb(ignore missing or invalid values). - Compute and print the average of one more numeric column (for example
runtime_minormetascore, ignoring missing values). - Count how many movies have
rating_imdb >= 8.0. - Report the time and space complexity for:
- first-match search task
- average computation task
Tip
You may find the following tips useful.
- Skip the header row using
next(reader)before processing the data. - Convert numeric values using
float()orint()before calculations. - Use a counter and a running total to compute averages (
total / count). - Use
breakto stop the loop once the first matching row is found.
- Do not use
pandas. - Handle invalid/missing numeric values safely.
- Keep your code readable with clear variable names.
## Homework (Session 1)
- File: `solutions/exercise-01-homework.py`
- Status: completed
- Notes:
- computed averages for rating and one additional numeric column
- handled missing or invalid values safely during computationsCreate a public GitHub repository for your homework. It is recommended to use one repository for all weekly submissions, for example: bda-homeworks.
Submit your work by sharing your repository link in the MS Teams channel
(Discussion forum for this class): MS Teams discussion forum.
This allows Stelios and the rest of the class to view and discuss your work 𐦂𖨆𐀪𖠋𐀪𐀪.