Posted on: 7th Oct 2026

DSM020 Data Programming in Python Coursework: Exploratory Data Analysis | University of London

DSM020 Data Programming in Python Coursework: Exploratory Data Analysis

Programme: MSc Data Science, University of London

Module: Data Programming in Python (DSM020)

Weighting: 30% of the total module grade

Coursework Description

Produce a proposal for a data science project of your choosing, which you intend to pursue in the second half of the module. Define the aims and objectives of your project, and use a range of programming techniques to make sure your data is suitable for analysis.

Use critical and analytical skills to explore your dataset through exploratory data analysis (EDA) and identify the key challenges of working with the data. This is your first substantial data processing pipeline and prepares your data for the final coursework assignment of the module.

Describe the formal approach you have taken, including design decisions along the way, such as how the data was captured or retrieved, summary statistics, and refined discussion that helps form concrete research questions for your second coursework assignment.

Submission: a single Jupyter Notebook and any related scripts or SQL files in a single ZIP archive. The notebook should describe your approach and include all processing used to manipulate, cleanse and sanitise the data, plus visualisations and tables. If your dataset exceeds 10MB, include a working sample.

Example Themes

You do not have to choose one of these; they are examples only.

Theme Scope Example project
Premier League Football English Premier League data from 20 February 1992 to the current season “How to get relegated: an analysis of poorly performing teams in the English Premier League”
Literary Masterpieces Famous plays, sonnets and poems “What’s in a name? An investigation into the names and content of the works of Shakespeare”
HTML and Markup Markup from one or more of the top 50 websites “An analysis of the semantic features of streaming websites”

Deliverables

Design a manageable data science project and acquire the necessary dataset in a usable form. Your notebook should include all acquisition steps, pre-processing and changes to the data, with a clear design rationale throughout. For this exercise you should:

  • Acquire and prepare your own dataset, making sure you are allowed to share it and anonymising any sensitive information.
  • Collate and/or manipulate the data into a usable format, possibly combining multiple sources to verify integrity and accuracy.
  • Create a database or flat-file format where needed, for example for very large datasets or frequent CRUD operations.
  • Where working with sensitive data, write Python that produces a realistic dummy dataset with sensible distributions and noise.
  • Explain the programming techniques used to prepare the data, including any command-line or SQL programming.
  • Outline the idea behind your project (context, significance, expected outcomes) and briefly detail what you intend to do with the data.
  • Consider any weaknesses or potential caveats in your approach.

Report Structure

Present your report in a single notebook that includes:

  • Introduction and context
  • A brief description of the dataset (or a sample output), including how it was obtained
  • High-level descriptions of the data, its source and its appropriateness for your path of exploration
  • A summary of key findings and insights
  • Discussion and critical analysis
  • Conclusion and further work
  • References to any resources used

Visualisations can be shown inline or exported separately (e.g. PNGs) if too large. Include a working sample of your dataset (maximum 10MB) and a requirements.txt file with instructions on how to replicate your approach.

Marking Rubric (10 marks each)

Element Example considerations
(a) Data is interesting enough to facilitate insights Sufficient complexity to demonstrate coding, data manipulation and analysis in Python.
(b) Data is relevant and the source is justified Origin described, reasons for selection, columns linked to research questions, format suitable for analysis.
(c) Project background is clearly defined Why the field matters, limitations, how the project addresses them, and fitness of the dataset.
(d) Dataset sufficiently prepared for analysis Illegal values removed, nulls handled, out-of-bound values checked, correct format (e.g. tokenised text).
(e) Ethics of data use considered Open or proprietary source, ownership of derivative data, risk of harm or discrimination, anonymisation.
(f) Clear rationale for data modifications Modifications are justified, add value and use advanced techniques where appropriate.
(g) Code is clean Functions for repeated processes, comments or markdown, clear and orderly code.
(h) Code is functional and reproducible Runs without errors, uses relative paths, explains libraries, avoids long wait times.
(i) Data captured using a technique Web scraping, database import/export, or justification for using a dataset as is.
(j) Readability of code Logical, systematic steps separated into cells, with an appropriate level of discussion.

Get Expert Help With Your DSM020 Data Programming in Python Coursework

Get Help By Expert

Not sure how to choose a dataset complex enough to score well, or how to show web scraping, cleaning and ethics clearly in one Jupyter Notebook? Our my assignment help data science experts support students with the DSM020 Data Programming in Python exploratory data analysis coursework, from framing a project proposal and acquiring data to building a reproducible pandas pipeline and writing clear EDA findings against all ten rubric elements. You can also review our computer and IT assignment samples, get help with the write-up through our report writing service, or browse more University of London assignment questions. If you are short on time, our do my assignment service gives personalised guidance so your notebook stays original.

Tags:- DSM020 UOL
Answer
AVM345 Airline Operations and Planning End-of-Course Assessment 2026
No Need To Pay Extra
  • Turnitin Report

    $10.00
  • Proofreading and Editing

    $9.00
    Per Page
  • Consultation with Expert

    $35.00
    Per Hour
  • Live Session 1-on-1

    $40.00
    Per 30 min.
  • Quality Check

    $25.00
  • Total
    Free

New Special Offer

Get 30% Off

Hire an Assignment Helper and Earn A+ Grade

UP TO 15 % DISCOUNT

Get Your Assignment Completed At Lower Prices

Plagiarism Free Solutions
100% Original Work
24*7 Online Assistance
Native PhD Experts
Hire a Writer Now

Facing Issues with Assignments? Talk to Our Experts Now! Download Our App Now!