Open to internships

Jayden Lee

Machine learning · UC Berkeley, Data Science

ML systems for messy, real-world data: fan behavior at Cal Athletics, transit-aware engagement at Oakland Roots, proteomics at a national lab.

Jayden Lee

I'm a Data Science student at UC Berkeley, with a domain emphasis in applied mathematics, who builds machine learning systems for data that doesn't behave: behavioral logs, mass spectrometry output, clinical time series.

Most of my work has been the unglamorous half of ML: getting 300,000 files through a pipeline without it falling over, deciding what a false discovery rate should be, lining up 36 spectral files that three tools each name differently. I like problems where the modeling is only interesting once the data engineering is honest.

Study
B.A. Data Science, applied mathematics emphasis, UC Berkeley
Degree
Expected May 2029
GPA
3.7
Group
Sports Analytics Group at Berkeley · Head of Business
Languages
Python · R · Java · TypeScript · SQL
Modeling
PyTorch · TensorFlow · scikit-learn · NumPy · Pandas · Seaborn · Matplotlib
Toolchain
Git · Docker · Kubernetes · Flask · React · Supabase
Speaks
English · Korean
Mail
jeyl@berkeley.edu
Code
github.com/jynlee7

1st place · 7th Annual Datathon at Berkeley

Clinical Intervention Simulator

A Monte Carlo framework that models 300 patient deterioration scenarios, feeding a Random Forest that optimizes interventions in real time.

Screenshot · 21:9

The simulation is vectorized with NumPy so a full sweep runs in seconds rather than minutes, which is what made it usable as an interactive tool instead of a batch job. It ships as a Streamlit app over a Pandas ETL pipeline that reconciles multimodal health data into one frame.

Built at the 7th Annual Datathon at Berkeley, where it took first place against more than 50 teams and over 200 participants.

Stack
Python · scikit-learn · NumPy · Pandas · Streamlit · Tableau
Scale
300 patient scenarios
App
Streamlit app on Deepnote

ChessBlitz

An open-source teaching platform in use at Berkeley Chess School.

Screenshot · 16:10

Flask and Supabase behind the storage layer, OpenRouter generating hints tailored to the position a student is actually stuck on, and a React dashboard where a teacher runs their own classroom: rosters, assignments, and who is struggling with what.

Stack
Flask · Supabase · OpenRouter · React · TypeScript · SQL
Status
Open source, in classroom use
Source
gitlab.com/devahschaefers/chessblitz

This Website

The machine you are reading this in. A desktop OS in vanilla TypeScript — no framework, no UI library, no CSS framework — built as a study in running AI agents on a codebase without losing control of it.

The build is multi-agent. Each task routes to a purpose-built skill — design, code review, QA — and agents run in isolated git worktrees so several can work at once without colliding in a merge. That only holds if the agents share a source of truth, so the rules live in machine-readable specs beside the code: what the product is, what the visual contract allows, what the game may and may not do. An agent that has read them cannot argue the interface into something else three sessions later.

Every change is gated behind a TypeScript build and an automated accessibility pass before it can land, and each release goes into a Keep a Changelog log under semantic versioning. The result is a codebase in the thousands of lines that has stayed consistent across dozens of agent sessions — which is the actual difficulty, and the reason this is on the list at all.

The wallpaper is a contour plot of a two-minimum loss surface. The surface function has one definition, imported by both the build script that writes the static SVG and the canvas that deforms under your pointer, so the two can never disagree about the shape of the field.

Stack
TypeScript · Vite · Claude Code · Git
Dependencies
2, both build-time
Source
github.com/jynlee7

Experience

  1. Feb – May 2026

    Data Analyst

    Cal Athletics · contract

    A pipeline through Snowflake and Colab that quantified purchases and fed a dynamic fan scoring system, plus clustering and supervised models predicting ticket-purchase propensity from behavioral data.

    Takeaway Understanding the data fully accelerates the work done after.

  2. Sep – Dec 2025

    Business Analyst

    Oakland Roots Sports Club · contract

    Built a fan-engagement scoring system in Python that folded real-time BART transit data into a marketing dashboard, and did the exploratory work behind it in Matplotlib and Seaborn.

    Takeaway The mentor won't report to you — you ask, and you report to them.

  3. Jun 2024 – May 2025

    Modeling & Simulation Intern

    Pacific Northwest National Laboratory

    A target-decoy statistical framework modeling false discovery rates for peptide identification. Deployed the GPU-accelerated proteomics workflows on the Tahoma HPC cluster under Slurm, configuring CUDA and batch processing so identification validation scaled across large mass-spectrometry datasets.

    Takeaway Do not assume a result is correct because it arrived. Ask the statistics whether it is wrong.

  4. Jun – Aug 2023

    Computational Biology Intern

    Pacific Northwest National Laboratory

    Built the file mapping and parsing pipeline that aligned 36 MGF spectral files across tool formats that each named them differently, then pushed 300,000+ mass-spec files through GPU-accelerated Casanovo inference on TensorFlow.

    Takeaway I like GitHub, and I want to do more open-source work.

B.A. Data Science, applied mathematics emphasis, UC Berkeley · GPA 3.7 · expected May 2029

Sports Analytics Group at Berkeley — Head of Business

Coursework: Optimization Models in Engineering · Data Structures and Algorithms · Data Engineering · Multivariable Calculus · Linear Algebra and Differential Equations

Your browser won't preview PDFs inline.

Download résumé
Kind
PDF document
Size
185 KB
Pages
1
Download

spots.map

A few blocks of Berkeley, and what I ate on them. Walk it and the places open one at a time.

Draft. Not finished, and visible so it cannot be forgotten.

Places I actually went

Twelve or so places in one Berkeley neighborhood, and what I ate at each. Walk the map and they open one at a time.

No places written up yet — the map below shows the storefronts standing empty, which is the truth of it.

Every place on this map is somewhere I actually went, written up in one file I edit by hand. Nothing here is generated and there are no filler entries — where I have not written anything, the storefront stands empty and the map says so rather than inventing a restaurant to fill the gap. That constraint is most of why it is worth walking.

It is a map rather than a list because a list gets skimmed. Twelve names and twelve one-line reviews are four seconds of scrolling and nothing retained. Making you walk two blocks to the next one buys each place the few seconds it takes to actually read it, which is the only thing I wanted from the format.

There is no score, no timer and no way to lose, and it does not ask you to come back tomorrow. Walking is one keypress per block, the map redraws only when you find something, and nothing on this page animates while you are standing still.