Back to the homepage

The program

Twelve weeks to statistical packages built from scratch

This is the first of the four phases. Three months to build the statistical foundation our machine learning work and everything after it will stand on.

How it works

  • We read a selected canonical book.
  • We convert the book to code.
  • We supplement everything with research papers.
  • We write notes and draw tldraw diagrams to explain everything. This is also how we know you understood things, even if you use AI.

The growing method

Twelve weeks, three phases. Everyone builds the same phase at the same time, then turns their own packages loose on new questions.

Phase 1 · Weeks 1-3

Descriptive statistics

Numeric-data validation, count, min, max, range, mean, median, mode, quantiles, IQR, sample variance and standard deviation, coefficient of variation, skewness, kurtosis, frequency tables, histogram bin counts, covariance and Pearson correlation, and stable mean and variance where it changes the result for real data.

  • Descriptive stats
  • Weeks 1-3
  • Pure Python
Phase 2 · Weeks 4-9

Probability distributions

One documented distribution interface: pmf or pdf, cdf, sf, ppf, mean, var, and reproducible random sampling. Bernoulli, Binomial, Geometric, Negative Binomial, Hypergeometric, Poisson, Uniform, Normal, Exponential, Gamma, Chi-square, Student t, and F. Explicit support, parameterization, invalid-parameter errors, and a tail-accuracy policy for every distribution. It does not become a general library for probability puzzles, set operations, or Bayes calculations.

  • Distributions
  • Weeks 4-9
  • Tested edge to edge
Phase 3 · Weeks 10-12

Inference and hypothesis testing

One-sample and two-sample confidence intervals for means. One-sample t test, Welch two-sample t test, paired t test, and one- and two-proportion tests. A result object with statistic, p-value, degrees of freedom, confidence interval, alternative, method, and assumptions. Condition checks and plain-language warnings: a function must not return a p-value when its own input rules fail.

  • Inference
  • Weeks 10-12
  • Assumptions checked

The book

We use the first half, up to and including chapter 8, of Introduction to Probability and Statistics for Engineers and Scientists, sixth edition, by Sheldon M. Ross. A clear, applied, upper-undergraduate book. It begins with data collection and descriptive statistics, moves through random variables and named distributions, then builds sampling distributions, estimation, and hypothesis tests. The later chapters carry on to regression, ANOVA, categorical analysis, resampling, and machine learning.

Our language of implementation is Python.

How the team works

Every person works on the same phase at the same time.

  1. 01Everyone works on descriptive statistics for weeks 1 through 3.
  2. 02Everyone works on probability distributions for weeks 4 through 9.
  3. 03Everyone works on inference and hypothesis testing for weeks 10 through 12.
  4. 04Everyone thinks about and works toward a novel individual research that uses the packages we develop. The paper, with related code and materials, is submitted sometime after week 12.

There are no permanent descriptive, probability, or inference teams.

Example weekly rhythm

DayWork
MondayA 45-minute planning call. Assign pages and material, exercises from the selected book, reviewers, and the rest.
Tuesday and WednesdayRead, solve, write notes, and start implementation.
ThursdayEvery contributor opens a real PR. It may still be a draft, but it must contain actual work.
FridayReview, test, and cut work that cannot meet the quality gate. Tldraw canvas explanation. PR and peer review deadline.
SaturdayOne-to-one meetings, in person or remote.
SundayCore team decides on progress and next steps.

Definition of done

Every feature PR must include:

  • Inputs, outputs, parameterization, support, assumptions, errors report, and examples.
  • A public-vault note linked to the textbook section, a tldraw diagram built with the tldraw Obsidian plugin, and the deeper source material if one was used.
  • Hand-worked results, usual cases demonstration, boundary cases, and tests.
  • A controlled comparison against a trusted library.
  • Review by somebody other than the author.
  • Proof you reviewed the work of other people as well. We cannot stress how important this is.
  • One documentation example that runs in a clean environment.

Release scope

Weeks 1-3

Descriptive statistics

  • Numeric-data validation and one clear missing-value policy.
  • Count, min, max, range, mean, median, mode, quantiles, IQR, sample variance, sample standard deviation, coefficient of variation, skewness, and kurtosis.
  • Frequency tables and histogram bin counts.
  • Covariance and Pearson correlation.
  • Stable mean and variance where it changes the result for real data.
  • Any additional material from the book.
Weeks 4-9

Probability distributions

  • This package models and evaluates distributions. It does not become a general library for solving arbitrary probability puzzles, set operations, counting rules, or Bayes calculations.
  • One documented distribution interface: pmf or pdf, cdf, sf, ppf, mean, var, and reproducible rvs where appropriate.
  • Bernoulli, Binomial, Geometric, Negative Binomial, Hypergeometric, Poisson, Uniform, Normal, Exponential, Gamma, Chi-square, Student t, and F.
  • Explicit support, parameterization, invalid-parameter errors, and a tail-accuracy policy for every supported distribution.
  • Sampling-distribution helpers required by the inference package.
  • Any additional material from the book.
Weeks 10-12

Inference and hypothesis testing

  • One-sample and two-sample confidence intervals for means.
  • One-sample t test, Welch two-sample t test, paired t test, and one- and two-proportion tests.
  • Chi-square goodness-of-fit and independence tests only if the team finishes the required work by the end of week 11.
  • A result object with statistic, p-value, degrees of freedom where relevant, confidence interval where relevant, alternative, method, and assumptions.
  • Condition checks and plain-language warnings. A function must not return a p-value when its own input rules fail.
  • Any additional material from the book.

The first 12 weeks

This schedule is flexible and can change as we go deeper.

WeekShared readingShared phase and build targetSaturday PR evidenceSunday presentation
1Ch. 1, introduction to statistics, data collection, samples and populations.Descriptive. Package skeleton, numeric input contract, missing-data decision, count, min, max, range.Empty input, invalid values, and a hand-checked small dataset. Book examples and exercises solved with our package.Explain the data contract.
2Ch. 2.1 to 2.4, describing and summarizing datasets.Descriptive. Mean, median, mode, quantiles, IQR, variance, standard deviation, frequency tables, histogram counts.Textbook examples reproduced by hand and in code. Book examples and exercises solved with our package.Show how mean and median disagree on a skewed dataset.
3Ch. 2.5 to 2.6, normal datasets, paired data, correlation.Descriptive. Covariance, Pearson correlation, skewness, kurtosis, stable online moments. Descriptive v0.1 release candidate.Constant values, tied values, numerical-stability comparison, documentation example. Book examples and exercises solved with our package.Demonstrate one summary that can mislead.
4Ch. 3, elements of probability.Distributions. Read probability as background only. Define the distribution interface and implement Bernoulli and Binomial.PMF sums to one, support and parameter errors, known moments. Book examples and exercises solved with our package.Explain the interface and what the package will not do.
5Ch. 4, random variables and expectation.Distributions. Implement Geometric, Negative Binomial, Hypergeometric, Poisson, plus moments and random sampling.Known values, seeded draws, CDF limits. Book examples and exercises solved with our package.Present one distribution and its parameterization.
6Ch. 5.1 to 5.4, special random variables.Distributions. Implement Uniform and Normal. Add shared cdf, sf, and ppf test cases.CDF and quantile round trips, tail values, comparison tolerance.Explain why tail probabilities are harder than ordinary values.
7Ch. 5.5 to 5.9, Normal, Exponential, Gamma, related distributions.Distributions. Implement Exponential, Gamma, Chi-square, Student t, and F.Boundary values, moments, extreme inputs, reference checks. Book examples and exercises solved with our package.Explain one numerical method or approximation.
8Ch. 6.1 to 6.3, sampling statistics and the central limit theorem.Distributions. Add sampling-distribution examples and the helpers inference will need.A reproducible CLT simulation and a written interpretation.Show the bridge between distribution code and inference.
9Ch. 6.4 to 6.6, sample variance and normal-population sampling.Distributions. Finish parameter validation, API consistency, examples, and release notes. Distributions v0.1 release candidate.Clean install and full distribution regression suite. Book examples and exercises solved with our package.Live review against a reference library.
10Ch. 7, parameter estimation and interval estimates.Inference. Implement result object, one-sample mean interval and test, and two-mean intervals.Hand-calculated textbook case and invalid-condition tests. Book examples and exercises solved with our package.Explain confidence intervals without the usual false claim.
11Ch. 8.1 to 8.4, significance levels and mean tests.Inference. Implement Welch two-sample t test, paired t test, and one- and two-proportion tests.Known p-values for each alternative, paired-data cases. Book examples and exercises solved with our package.Explain Type I error, Type II error, and power with one dataset.
12Ch. 8.5 to 8.7, variance, Bernoulli, and Poisson tests.Inference. Finish tests, documentation, mini-study, changelog, release tag, and retrospective. Start chi-square only if all required tests passed by Friday.Clean install, reproducible mini-study, all examples and links checked. Book examples and exercises solved with our package.Final release. Each person explains one contribution and one current limitation.

How to read and research

The book is mandatory. Each person reads the assigned pages and solves exercises from the book with the functions they developed. Implementation is supported by the book, a minimum of two canonical research papers of importance in statistics, evidence of the same results using numpy, pandas, statsmodels, or scipy, a tldraw canvas explanation for presentation, and a note in the open Eskolx Obsidian vault linking all of it together.

Example papers include Welford for online variance, a numerical reference for Normal CDF and quantiles, or Welch for unequal-variance testing.

At the end of the week, each person reviews the work of a minimum of two other people and verifies it is correct.

The program starts at the repo.

Want to join phase one? Write to eskolxlabs@gmail.com and tell us a little about yourself.

Eskolx Labs

Learn deep, build expertise.

Contacteskolxlabs@gmail.com

Statistical libraries, rebuilt from scratch, in the open.

Back to the top

© 2026 Eskolx Labs · MIT License