Rust data engineer. Solana and on-chain data infrastructure.
I build streaming data infrastructure in Rust. Most of my work sits where a firehose of on-chain state meets a database that has to answer questions about it: gRPC subscriptions, Anchor account decoding, columnar storage, and the plumbing that keeps a pipeline running unattended for weeks.
klend-indexer is the longest running piece of that. It subscribes to Kamino Lend account writes over a Yellowstone gRPC stream, decodes them with Anchor discriminators, and lands them in ClickHouse. On GCP it accumulated roughly 1.1M account updates and 180K decoded obligation snapshots. It's frozen now, on purpose: Yellowstone can't re-serve old slots, so the history was exported to Parquet before the stream was stopped. Designing around that one constraint taught me more than the rest of the pipeline did.
karat67 came out of running that
thing. An indexer can pass every freshness and throughput check while writing
empty or malformed rows, and every dashboard built on top inherits the error.
karat67 checks indexed data against the chain itself: account shape against
the IDL, byte-level reconciliation of sampled accounts, and slot completeness
against getBlocks (skipped slots aren't gaps, which is the part people get
wrong). Freshness and liveness are deliberately out of scope. Every indexer
already has those.
I also do network measurement research at Oregon State's ASTRO Lab under Dr. Zane Ma, on how much validator identity leaks at Solana's gossip layer and how that leak is weighted by stake and concentrated by cloud provider.
Solana is the domain I know best. What sits underneath it is Rust and data engineering, and that part travels.
Off the clock I argue about the conceptual mechanics of Devil Fruits in One Piece.
- karat67: integrity checks for
Solana indexers, as a Rust crate and a
karatCLI. Three checks (shape, reconciliation, completeness), JSON output, built to drop into CI next to an existing pipeline. This is the active one. - klend-indexer: Rust to Yellowstone gRPC to ClickHouse. Account-level decode, slot checkpointing, gap detection, lag metrics, and a freeze/resume runbook. Frozen, maintained, the dataset lives in Parquet on GCS.
- S-NodeFinder: gossip-layer measurement at OSU ASTRO Lab. Collection is running, analysis is paused.
- Cold-path storage: Parquet and DataFusion beside ClickHouse, so one dataset serves live queries and cheap historical scans. Planned, not built.
Happy to talk streaming data, Rust, or Solana internals.
"The clock is the dataset: a stream you stop observing is history you cannot get back."


