# Synthetic data analysis example: study hours and a fictional quiz score

> **Teaching example only.** Every row in the companion CSV was generated for this exercise. These values were not collected from students, courses or an institution. They do not support claims about real learners, learning outcomes or cause and effect.

Companion file: [`synthetic-study-hours-data.csv`](./synthetic-study-hours-data.csv)

## Data dictionary and provenance

| Column | Meaning | Type |
| --- | --- | --- |
| `record_id` | Invented row label such as SYN001; it is not a person or record key from a real system | Text |
| `study_hours` | Fictional number of study hours used to create an example row | Decimal |
| `practice_quizzes` | Fictional count used to create an example row | Integer |
| `fictional_quiz_score` | Invented score from 0 to 100; not an observed assessment result | Integer |

The 40 records are deterministic teaching values, created by a simple formula with a fixed pattern of variation. The synthetic values intentionally have positive correlations so that learners can practise describing an association. Because this is generated data, that pattern is a feature of the construction, not a discovery about an educational setting.

## Reproduce descriptive summaries

Save this code as `analyse.py` beside the CSV and run `python analyse.py`. It uses only the Python standard library.

```python
import csv
import math
from pathlib import Path

path = Path(__file__).with_name("synthetic-study-hours-data.csv")
with path.open(newline="", encoding="utf-8") as source:
    rows = list(csv.DictReader(source))

hours = [float(row["study_hours"]) for row in rows]
quizzes = [int(row["practice_quizzes"]) for row in rows]
scores = [int(row["fictional_quiz_score"]) for row in rows]

def mean(values):
    return sum(values) / len(values)

def pearson_r(left, right):
    left_mean, right_mean = mean(left), mean(right)
    numerator = sum((x - left_mean) * (y - right_mean) for x, y in zip(left, right))
    left_ss = sum((x - left_mean) ** 2 for x in left)
    right_ss = sum((y - right_mean) ** 2 for y in right)
    return numerator / math.sqrt(left_ss * right_ss)

print(f"Rows: {len(rows)}")
print(f"Mean study_hours: {mean(hours):.3f}")
print(f"Mean practice_quizzes: {mean(quizzes):.3f}")
print(f"Mean fictional_quiz_score: {mean(scores):.3f}")
print(f"Score range: {min(scores)} to {max(scores)}")
print(f"Pearson r (study_hours, fictional_quiz_score): {pearson_r(hours, scores):.3f}")
print(f"Pearson r (practice_quizzes, fictional_quiz_score): {pearson_r(quizzes, scores):.3f}")
```

Expected output for the file as supplied:

```text
Rows: 40
Mean study_hours: 2.255
Mean practice_quizzes: 3.500
Mean fictional_quiz_score: 65.975
Score range: 47 to 86
Pearson r (study_hours, fictional_quiz_score): 0.758
Pearson r (practice_quizzes, fictional_quiz_score): 0.515
```

## What the numbers mean—and do not mean

The values summarize only these 40 generated rows. The two Pearson coefficients describe linear association inside this constructed example; they do not test a research hypothesis, quantify the probability that a hypothesis is true, establish statistical significance, or show that study time or quiz practice caused a score. No p-value or confidence interval is included because the data are synthetic and the purpose is to practise transparent descriptive reporting, not population inference.

For an actual project, first establish a suitable question, data provenance, measurement definitions, sampling and ethical permissions. Check that the analysis fits the design and assumptions, document cleaning and exclusions, report uncertainty where appropriate, and interpret the result within its limits. Consult your supervisor or a qualified methods instructor about the analysis your programme requires.

## Reproducibility checklist

- Keep the original CSV unchanged and save any cleaning steps separately.
- Record the script, software and version used to produce a result.
- Label any table or chart made from this file as **synthetic teaching data**.
- Never copy these outputs into a report as findings about real people or institutions.
- For real participant data, follow approval, consent, access and data-protection requirements.

Further reading: [ASA Statement on Statistical Significance and P-Values](https://www.amstat.org/asa/files/pdfs/P-ValueStatement.pdf) and [OpenLearn: Getting started with statistics](https://www.open.edu/openlearn/mod/oucontent/view.php?id=106071&section=4.2).
