Data Platforms · Open Source · Analytics
Data
I build data platforms for specialized fields, turning domain knowledge into reliable software, pipelines, and analytical tools.
- Applied Science
- Digital Assets
- Research Infrastructure
4,035 GitHub contributions from the past year, visualized as data flowing through a pipeline.
Roles
Current Work
- Data Platform EngineeratHypertrial
Building open-source data infrastructure that makes complex datasets easier to use.
- Senior Data ScientistatFigment
Leading analytics engineering for digital-asset staking with $15B+ under stake.
Technical LeadatTrilemma FoundationLeading data platform work across research and product, connecting analytical infrastructure with open-source goals.
Upstream
Open-Source Contributions
Selected fixes and improvements I've contributed to open-source data tools.
- rocky
Improved SQL transformation correctness: blocked unsafe raw dbt incrementals, aligned plan/apply model scope, and improved cast type inference.
- Rust
- dlt
Added JWT auth without required scopes, preserved paginator stop conditions, fixed filesystem file-extension handling, and tightened SCD2 merge-key matching and TOML config serialization.
- Python
- Apache Iceberg
Fixed Parquet notNaN edge cases, SQL catalog connection-pool validation, truncate-transform ordering, and missing manifest partition summaries across Java, Rust, and Python.
- Java
- Rust
- Python
Apache PolarisCorrected CLI exit codes so table and setup failures are detectable by scripts and automation.
- Python
- OpenLineage
Added async HTTP event delivery in the Python client and WHERE/ARRAY subquery input lineage in the SQL parser.
- Python
- Rust
- dqx
Fixed null-safe result joins so custom SQL and grouped data-quality checks keep nullable-key violations instead of silently dropping them.
- Python
- Velox
Fixed an edge case that could invalidate resized dictionary vectors over empty bases in Meta's C++ execution engine.
- C++
- duckle
Fixed Python pipeline limit serialization so limit(n) reaches the engine instead of silently falling back to LIMIT 100.
- Python
- slayer
Masked inline BigQuery credentials in datasource output and hardened YAML storage identity handling.
- Python
Architecture