Open to 2026 internships

Jayden Lee

Machine learning · UC Berkeley, Data Science & Computer Science

ML systems for messy, real-world data: fan behavior at Cal Athletics, LLM training efficiency at MatX, proteomics at a national lab.

Draft. Not finished, and visible so it cannot be forgotten.

How I work

Replace this lede with two or three sentences on what you actually believe about building software. Not what sounds good in an interview — what you would still say after a bad week.

Jayden Lee

I'm a Data Science and Computer Science student at UC Berkeley who builds machine learning systems for data that doesn't behave: behavioral logs, mass spectrometry output, clinical time series.

Most of my work has been the unglamorous half of ML: getting 300,000 files through a pipeline without it falling over, deciding what a false discovery rate should be, arguing with hardware engineers about what "efficiency" means. I like problems where the modeling is only interesting once the data engineering is honest.

Study
B.A. Data Science and B.A. Computer Science, UC Berkeley
GPA
3.63
Group
Open Project · Sports Analytics Group at Berkeley
Mail
jeyl@berkeley.edu
Code
github.com/jynlee7

1st place · 7th Annual Datathon at Berkeley

Clinical Intervention Simulator

A Monte Carlo framework that models 300 patient deterioration scenarios, feeding a Random Forest that optimizes interventions in real time.

Screenshot · 21:9

The simulation is vectorized with NumPy so a full sweep runs in seconds rather than minutes, which is what made it usable as an interactive tool instead of a batch job. It ships as a Streamlit app over a Pandas ETL pipeline that reconciles multimodal health data into one frame.

Built at the 7th Annual Datathon at Berkeley, where it took first place against more than 50 teams and over 200 participants.

Stack
Python · scikit-learn · NumPy · Pandas · Streamlit · Tableau
Scale
300 patient scenarios
Source
github.com/jynlee7

ChessBlitz

An open-source teaching platform in use at Berkeley Chess School.

Screenshot · 16:10

Flask and Supabase behind the storage layer, OpenRouter generating hints tailored to the position a student is actually stuck on, and a React dashboard where a teacher runs their own classroom: rosters, assignments, and who is struggling with what.

Stack
Flask · Supabase · OpenRouter · React · TypeScript · SQL
Status
Open source, in classroom use
Source
github.com/jynlee7

Experience

  1. Feb 2026 – now

    Machine Learning Researcher

    Cal Athletics · contract

    Clustering and supervised models that predict ticket-purchase propensity from behavioral data, plus multivariate regression behind a dynamic fan scoring system. Outputs deploy into CRM workflows for segment-level personalization.

  2. Dec 2025 – now

    Machine Learning Consultant

    MatX · contract

    Leading eight engineers building a competitive benchmarking platform for a Series A AI-chip startup, evaluating LLM training efficiency across participant-submitted models. Docker containerization and Kubernetes job scheduling isolate each participant. Defined decode efficiency metrics directly with MatX hardware engineers.

  3. Sep – Dec 2025

    Business Analyst

    Oakland Roots Sports Club · contract

    Built a fan-engagement scoring system in Python that folded real-time BART transit data into a marketing dashboard.

  4. Jun 2024 – May 2025

    Modeling & Simulation Intern

    Pacific Northwest National Laboratory

    A target-decoy statistical framework modeling false discovery rates for peptide identification, with a precision-recall threshold optimization pipeline. Ran GPU-accelerated proteomics workflows on the Tahoma HPC cluster under Slurm.

  5. Jun – Aug 2023

    Computational Biology Intern

    Pacific Northwest National Laboratory

    Pushed 300,000+ mass-spec files through a GPU-accelerated Casanovo inference pipeline, and benchmarked 18 bacterial metaproteomics datasets with Pandas and Seaborn.

B.A. Data Science, UC Berkeley · GPA 3.7

Sports Analytics Group at Berkeley

Your browser won't preview PDFs inline.

Download résumé
Kind
PDF document
Size
181 KB
Pages
1
Download

spots.map

A few blocks of Berkeley, and what I ate on them. Walk it and the places open one at a time.

Draft. Not finished, and visible so it cannot be forgotten.

Places I actually went

Twelve or so places in one Berkeley neighborhood, and what I ate at each. Walk the map and they open one at a time.

No places written up yet — the map below shows the storefronts standing empty, which is the truth of it.

Every place on this map is somewhere I actually went, written up in one file I edit by hand. Nothing here is generated and there are no filler entries — where I have not written anything, the storefront stands empty and the map says so rather than inventing a restaurant to fill the gap. That constraint is most of why it is worth walking.

It is a map rather than a list because a list gets skimmed. Twelve names and twelve one-line reviews are four seconds of scrolling and nothing retained. Making you walk two blocks to the next one buys each place the few seconds it takes to actually read it, which is the only thing I wanted from the format.

There is no score, no timer and no way to lose, and it does not ask you to come back tomorrow. Walking is one keypress per block, the map redraws only when you find something, and nothing on this page animates while you are standing still.