Can the measurement be trusted?
Evaluation validity, uncertainty calibration, ML inference semantics, autograd correctness, and reusable scientific-ML infrastructure.
I'm Tsolmondorj Natsagdorj. I find where an abstraction stops matching reality, then build the test that exposes the gap.
My work spans AI evaluation, high-performance networking, scientific ML, protocol correctness, and autonomy. The recurring question is the same: what evidence would prove the system wrong?
Real RDMA execution found four defects after the synthetic suite was green—and rejected the first attempted fix. 90-second evidence view →
A prompt audit found severe provenance leakage in 13/24 examples, so I downgraded the original model score instead of defending it. 90-second evidence view →
Eight selected merged patches across MatGL, MLX-LM, Alloy, and uutils: correctness reviewed at somebody else's boundary. 90-second evidence view →
Evaluation validity, uncertainty calibration, ML inference semantics, autograd correctness, and reusable scientific-ML infrastructure.
RDMA device state, compatibility semantics, protocol invariants, test architecture, and autonomy work where targets stay targets until hardware produces evidence.
BRIDGE-bench's historical score is contaminated evidence after the 13/24 leakage audit. Current work is zero-cost measurement design, sanitizer invariants, and matched controls; there is no claimed sanitized rerun.
Aiur CARRIER-P0 reduces the carrier idea to one falsifiable interaction: mechanically verified recovery of one small aircraft. Physical recovery performance remains a target until telemetry exists.
Merged work spans MatGL uncertainty/autograd, MLX-LM sampling defaults, Alloy EIP-712 recursion, and uutils date/time compatibility.
Real Linux RDMA verbs and sysfs state exposed four defects after 197 unit tests passed. The first attempted correction still failed on-device.
The benchmark's strongest finding became a validity failure in the benchmark itself. The original score is retained as contaminated historical evidence, not clean capability performance.
The pretrained mean ranked useful candidates at roughly 5× the random screening rate on the tested proxy task; the tested MC-dropout uncertainty signal was anti-correlated with error.
Design, simulation, CAD, controller logic, and physical acceptance gates for an airborne recovery concept. Unlike the projects above, recovery performance is not yet observed.
Why these projects belong together → · public corrections → · all writing →