Kristian M. Eschenburg

Kristian M. Eschenburg

Senior Data Platform Engineer

Seattle, WA

About Me

I’m a Senior Data Platform Engineer at a biologics company in Seattle, where I build the data, platform, and ML systems that help scientists make use of our large-scale biomanufacturing data. My role has encompassed everything from R&D work building machine learning models for antibody design and property prediction, to cloud infrastructure design and setup, to building data orchestration pipelines.

Before industry, I did my PhD in biomedical engineering at the University of Washington, where I developed graph neural network algorithms for processing brain MRI images to better understand cortical architecture and its relationship to disease diagnosis and progression.

When I’m not working, I’m usually backcountry skiing, trail running, or out in the garden. I also love to travel and like practicing new languages.

Interests

  • Data platforms and orchestration
  • ML training, deployment, and inference infrastructure
  • Applied machine learning for scientific data

Education

  • PhD in Biomedical Engineering, 2021

    University of Washington

  • BSc, 2012

    University of California, Los Angeles

Experience

 
 
 
 
 

Just-Evotec Biologics

Apr 2022 – Present Seattle, WA
  1. Senior Data Platform Engineer
    Oct 2025 – Present
    Details
    • Modernized the Dagster orchestration platform on AWS ECS, decoupling platform infrastructure from pipeline code for independent deploys and horizontal auto-scaling. Terraform templates cut new-pipeline provisioning from hours to under 5 minutes. Operate 3 production pipelines spanning upstream process development, ERP, and finance.
    • Built and maintain a centralized schema registry: 65+ versioned, backwards-compatible data contracts across instrument ingestion, orchestration pipelines, and lakehouse writes, catching data quality issues at the point of entry before they reach downstream models and analytics.
    • Led data platform integration for a new automated mini-bioreactor platform, a capital investment to improve scale-up analyses, now in production: built the ingestion pipelines, linkage of mini-bioreactor runs to bench- and manufacturing-scale bioreactor data, and the analytical tooling for scale-up comparisons.
    • Own the data lifecycle for two SCADA production databases growing ~100 GB/month: AWS Backup snapshots, Step Functions exports of monthly Parquet snapshots to S3, Glacier tiering, and Athena for historical queries.
    • Architecting a regulatory-grade audit framework (data lineage, write-time schema versioning, user access logs) on immutable DynamoDB stores for on-demand FDA and DoD information requests.
    • Leading design of an experiment-management platform replacing 15+ legacy tools company-wide: a multi-service monorepo of FastAPI backends and a Next.js frontend over PostgreSQL, with shared Pydantic contracts, deployed via Terraform.
  2. Senior Data Scientist
    Apr 2022 – Oct 2025
    Details

    ML engineering

    • Designed and built in-house deep learning models for structure-based antibody sequence design (inverse folding): message-passing graph encoder-decoders over protein structure graphs, with autoregressive and parallel samplers for generating candidate sequence designs.
    • Built the distributed training stack (PyTorch DDP / torchrun on multi-GPU Slurm nodes, checkpoint/resume, rank-aggregated metrics, Hydra configs) and ran systematic experiments on masking strategy, model depth, and validation design.
    • Built dataset pipelines to pre-featurize Protein Data Bank complexes (DIPS, DB5, SAbDab) for faster training.
    • Took the model to production as a versioned Python package (unit tests, GitLab CI publishing to an internal registry), containerized and deployed on ECS Fargate.
    • Fine-tuned protein language models for sequence-liability and thermal stability prediction, deployed to production alongside the structure-based design models and versioned in MLflow.

    Data & platform engineering

    • Architected a serverless, event-driven AWS ingestion pipeline (S3, EventBridge, Lambda, ECS, SQS) for 7 scientific instruments with near-100% uptime, cutting time from raw assay file to analytics-ready data from days or weeks to under 30 minutes and closing a provenance gap where raw files sat on scientists’ laptops.
    • Introduced Terraform as the org-wide infrastructure-as-code standard and led migration of 35+ microservices from EC2 to ECS across test and production environments.
    • Built FastAPI data services, Dagster gold-layer pipelines, and Dash/Plotly dashboards used by 60+ scientists across 6 functional groups.
 
 
 
 
 

Curi Bio

Jun 2021 – Apr 2022 Seattle, WA
  1. Data Scientist
    Details
    • Developed deep learning models to predict stem cell differentiation outcomes from high-throughput microscopy images, reducing material resource costs by upwards of 25% for the associated research stage.
    • Built contractility waveform analysis pipelines for engineered cardiac and skeletal myocytes, using signal processing to characterize the impact of therapeutics on muscle cell function.
 
 
 
 
 

University of Washington, Integrated Brain Imaging Center

Sep 2014 – Jun 2021 Seattle, WA
  1. PhD Graduate Student, Biomedical Engineering
    Details
    • Built a turn-key orchestration pipeline for processing 1000+ functional and diffusion MRI scans (>1.5TB) on a GPU-backed HPC system.
    • Developed graph neural networks for human MRI segmentation and biomarker generation, improving accuracy 15%+ over standard CNNs and improving test-retest reliability of patient-specific segmentations by 6% across clinical scanning sessions.
    • Applied dynamic mode decomposition (DMD) to fMRI brain dynamics, outperforming state-of-the-art ICA at identifying canonical activation networks while requiring shorter scanning sessions.
    • Designed a spatial statistical modeling approach for analyzing variability in the topography of functional brain connectivity, with results aligning with long-standing theories of hierarchical brain organization.
    • Awarded a highly selective 3-year ARCS Washington Research Foundation fellowship.
 
 
 
 
 

Internships

Jun 2016 – Jun 2017 Washington State
  1. Software Engineering Intern, Phase Genomics
    Apr 2017 – Jun 2017
  2. Data Science Intern, Pacific Northwest National Laboratory
    Jun 2016 – Sep 2016

Skills and Technologies

Cloud & Infrastructure

  • Compute: ECS/Fargate, Lambda; Kubernetes (personal projects)
  • Events & storage: EventBridge, S3, DynamoDB
  • Access & observability: IAM, Cognito, CloudWatch
  • Infrastructure as code: Terraform

Data Engineering

  • Orchestration: Dagster
  • Lakehouse: Delta Lake, Parquet
  • Transactional stores: PostgreSQL
  • Governance: schema registries, versioned data contracts

ML & AI

  • Modeling: PyTorch (DDP), DGL, scikit-learn
  • Tracking & registry: MLflow
  • Managed LLMs: Bedrock
  • Applied: protein language models, graph neural networks

Applications & Delivery

  • APIs: FastAPI, Pydantic
  • Frontend: Next.js, React, TypeScript
  • Build & deploy: Docker, GitHub Actions CI/CD
  • Practices: test-driven development, monorepos

Recent Posts

Kubernetes for ML Engineering, Part 2: Isolation — Networking and RBAC

Isolation boundaries in a shared Kubernetes cluster, who can reach what on the network, and who can do what against the API server.

Kubernetes for ML Engineering, Part 1: Why, and Two Services Talking

Why I’m learning Kubernetes for ML engineering, and getting two services talking on a local cluster.

Correlation IDs and Request Lineage Across Services

When a scientist says a dashboard was slow this morning, our logs couldn’t tell us which requests were theirs. I built a request context, stored in a ContextVar and forwarded as HTTP headers, that stamps request, session, user, and service IDs onto every log line. The post covers the middleware, why a logging filter beats a LoggerAdapter, and fixes for noisy health checks and unhelpful 403s.

Service-to-Service Auth with Cognito Scopes

Why we replaced a shared bearer token with Cognito’s client-credentials flow, so each service has its own identity and each endpoint requires a specific scope. Covers resource servers, custom scopes, token validation, and a couple of ALB rule gotchas.

Deploying Dagster to AWS ECS, Part 2: Pipelines

The per-pipeline Terraform module that plugs into the platform, with two task definitions per pipeline, Cloud Map registration so the daemon can find each code server, four IAM roles, and secrets handling.

Research

Applied deep learning on brain imaging — the same methods I now apply to protein design.

Learning Cortical Parcellations Using Graph Neural Networks

We examine the utility of graph neural networks for the purpose of learning cortical segmentations. We show that attention-based transformer networks significantly outperform conventional GCN and linear feed-forward variants for the purpose of generating accurate reproducible cortical maps.

Linear Mapping of Cortico-Cortico Resting-State Functional Connectivity

Using non-linear dimensionality reduction of functional brain connectivity patterns, and multivariate spatial statistics to characterize the functional embeddings, we analyze the spatial relationships between pairs of cortical regions to better examine how pairs of cortical regions connect and relate to one another.

Extracting Reproducible Time-Resolved Resting State Networks Using Dynamic Mode Decomposition

In this paper, we develop a novel method based on dynamic mode decomposition (DMD) to extract resting-state networks from short windows of noisy, high-dimensional fMRI data, allowing RSNs from single scans to be resolved robustly at a temporal resolution of seconds. This automated DMD-based method is a powerful tool to characterize spatial and temporal structures of RSNs in individual subjects.

Automated Connectivity-Based Cortical Mapping Using Registration-Constrained Classification

In this analysis, we propose the use of a library of training brains to build a statistical model of the parcellated cortical surface to act as templates for mapping new MRI data.

Registering Cortical Surfaces Based on Whole-Brain Structural Connectivity and Continuous Connectivity Analysis

We present a framework for registering cortical surfaces based on tractography-informed structural connectivity. We define connectivity as a continuous kernel on the product space of the cortex, and develop a method for estimating this kernel from tractography fiber models.