Felipe Parodi

Independent ML consulting

Find out whether the model result is real—and what to test next.

I help technical founders, ML teams, and research groups audit evaluations, diagnose difficult model failures, and turn research ideas into reproducible prototypes.


Who I help

Teams with a concrete model, dataset, paper, prototype, or technical decision—not a vague request to “add AI.”

01

Technical founders

Pressure-test an evaluation or prototype before committing a larger build.

02

Applied ML teams

Turn an unstable metric or model failure into a ranked, testable diagnosis.

03

Research groups

Implement a paper or experimental idea with a baseline, evaluation, and usable handoff.

Services

Each engagement is bounded around one decision, failure, or method. Larger implementation work is scoped only after the first evidence is in.

01

Model Evaluation Audit

Review one AI/ML claim for leakage, confounds, weak controls, inappropriate metrics, and missing failure coverage.

  • Evaluation risk map
  • Prioritized experiment plan
  • Written audit and readout
02

ML Debugging Sprint

Reproduce one difficult failure, rank root causes, and run the smallest experiments that distinguish them.

  • Root-cause hypothesis tree
  • Targeted diagnostics
  • Reproducible findings and handoff
03

Research-to-Prototype

Turn one paper or technically specific idea into reproducible code and a decision-quality evaluation.

  • Baseline + agreed approach
  • PyTorch implementation
  • Results memo and next step

Selected work

Public work spanning evaluation design, model internals, computer vision, clinical video, and multimodal measurement.

Controlled experiments on DINO vision transformers showed that a standard intervention can overstate register-token dependence because the network compensates. Related work uses matched controls to audit visual-representation claims.

Computer vision · 2025

PrimateFace

A >500K-image, >60-genera toolkit for cross-species face analysis, packaged as an installable library with a labeling GUI, hosted demo, documentation, and tutorials.

A low-shot vision-language pipeline for classifying provider attention in live clinical video, with 91% zero-shot and more than 98% fine-tuned classification accuracy.

A synchronized 30-camera, 3D-tracking, audio, behavior, and neural-analysis platform built to recover reliable structure from naturalistic data.

How engagements work

A short path from an ambiguous technical concern to evidence the team can act on.

Define the decision

Agree on one claim, failure, or feasibility question and what would change as a result.

Inspect the evidence

Review the data, code, metrics, baselines, constraints, and known failure modes.

Run decisive tests

Prioritize the smallest experiments that separate the leading explanations.

Hand off clearly

Deliver reproducible artifacts, a short memo, and a direct readout of what to do next.

Felipe Parodi

About Felipe

I am a computational neuroscientist and machine-learning researcher with a Ph.D. from the University of Pennsylvania. My work includes first-author research on model evaluation and interpretability, an LLM evaluation system built at Google/YouTube, open-source computer-vision tools, and multimodal measurement systems for difficult real-world data.

I work primarily in Python, PyTorch, computer vision, multimodal ML, representation analysis, and experimental design.

View research

Have a result you do not fully trust?

Send a short description of the model, evidence, and decision you are facing. Please do not include confidential data, credentials, or sensitive records in the first message.

felipe@pipeparodi.com