Pedro Paiva
Data · Machine Learning · Product

Hi! I'm Pedro.

I turn messy data into models, tools, and products people can actually use.

I'm a Computational Sciences student at Minerva University working across data science, machine learning, statistical modeling, and applied AI. I like projects that start with a real question and end with something interpretable, useful, or deployable.

Open to 2027 opportunities Python · R · SQL ML · Statistics · Data Products
move your cursor anywhere to paint

selected work

Featured work

A few projects where the analysis, engineering, and final product all matter.

Spotify skip prediction project visualization
Machine learning study

Predicting Spotify Skips from Song Intros

I used 92,445 personal listening events to test whether a listener's decision to skip a track is better predicted from the first seconds of audio or from behavioral context.

92,445 listening events
61.8% audio CV accuracy
75.1% context CV accuracy
PyTorch AST CNN Behavioral Data
Bayesian football attendance modeling results
Bayesian statistics

Modeling Sports Attendance in Argentina

Hierarchical Bayesian modeling of Argentine football attendance using team effects, weekday effects, Negative Binomial likelihoods, posterior predictive checks, and missing-data estimation.

PyMC Bayesian Statistics MCMC PSIS-LOO

Technical work

Smaller projects where I explored algorithms, simulation, optimization, and applied AI.

Simulation

Wildfire Spread on a Real Deforestation Frontier

Cellular automata simulation built around real Amazon deforestation patterns to test how targeted interventions change simulated fire spread.

Result: targeted interventions reduced simulated burned area by up to 83.8% across 500 Monte Carlo runs.
Applied AI

RememberMe: AI Memory Support

A 24-hour hackathon prototype designed to make personal videos, audio, and notes searchable so relevant context can be surfaced when someone needs help recalling a memory.

UC Berkeley Hack for Impact: Best Hack for Quality of Life and Best Use of TwelveLabs.
Algorithms

Decoding Relationships Between Genes

Genealogical tree reconstruction from DNA sequences using Longest Common Subsequence similarity, greedy construction, and global dynamic programming.

Optimization: global dynamic programming improved the objective score from 1336 to 1442 compared with the greedy approach.
Data Structures

Constraint-Aware Task Scheduler

Object-oriented Python scheduler using a max-heap to prioritize tasks while accounting for dependencies, urgency, deadlines, duration, and time windows.

Stress testing: evaluated scheduling behavior with workloads of up to 1,000 tasks and analyzed the main O(n²) bottleneck.

Experience

Teaching, research, community building, and products for real users.

AURA Application Mentorship

Co-founder · Mentor
Brazil

Co-founded an education initiative supporting students applying internationally. Built resources, digital infrastructure, classes, advising systems, and application workflows used across multiple cohorts.

Yale Young Global Scholars

Instructor
Yale

Designed and taught a seminar on how algorithmic systems shape what people see, hear, and click, using active-learning methods with international high-school students.

MateMarote · UTDT

Research Assistant
Buenos Aires

Worked with long-running behavioral research data, digitizing historical records and building R and Python workflows for schema design, validation, quality assurance, and analysis.

Minerva University

Teaching Assistant
Global

Supported students across computational and analytical coursework, including formal analysis, quantitative reasoning, and interdisciplinary problem solving.

Writing

Notes from the intersection of data, technology, history, and moving around.

Out of Distribution

Notes from places, datasets, and everything that doesn't quite fit.

Out of Distribution is where I write about the things I keep encountering outside of my technical work: history, cities, technology, data, culture, travel, and the strange ways they end up overlapping.

View over Seoul accompanying the latest Out of Distribution article
Latest

I keep ending up out of distribution

Korea, history, borders, datasets, and what happens when the categories we use to describe the world stop fitting quite as neatly as they first appeared.

About

I'm Pedro, a Computational Sciences student at Minerva University. Most of my work sits somewhere between machine learning, statistics, product thinking, and data engineering.

I'm especially interested in projects where modeling choices have consequences outside the notebook: public-health forecasting, behavioral prediction, decision systems, recommendation, human-computer interaction, and data products that people can actually explore.

Outside of technical work, I teach, mentor students applying internationally, write, study languages, and spend a lot of time thinking about how technology behaves once it leaves the model and meets actual humans.

Python SQL R PyTorch scikit-learn PyMC pandas NumPy Statistics Machine Learning Forecasting Data Products

Let's build something useful.

I'm always interested in conversations about data, machine learning, product, research, and opportunities where technical work meets real-world problems.