Atlas AI
Distributed AI infrastructure platform for transformer systems, inference optimization, scaling analysis, observability, and performance engineering.
ML systems projects spanning infrastructure, performance engineering, distributed training, search evaluation, and reproducible tooling.
Distributed AI infrastructure platform for transformer systems, inference optimization, scaling analysis, observability, and performance engineering.
Analyzes LeRobot datasets through URDF-based forward kinematics, dual-arm trajectory playback, and interactive 3D voxelized workspace-coverage visualization.
Evidence-backed performance repair system that detects benchmark regressions, creates bounded Codex repair goals, verifies fixes through protected benchmark evidence, and preserves human approval before merge.
Search evaluation and regression-detection platform for benchmarking TF-IDF, BM25, and hybrid retrieval systems across relevance, latency, and query-level failure behaviour.
Systems-oriented profiler for analyzing communication overhead, memory bottlenecks, scaling efficiency, and distributed training behavior in large-scale ML workloads.
Distributed training simulator analyzing scaling efficiency, communication overhead, and system-level bottlenecks across data-parallel workloads.
Reverse-mode autodiff engine with dynamic computation graphs and topological backpropagation. Verified gradient correctness and analyzed trade-offs between memory usage, execution efficiency, and graph flexibility.
CLI-based ML reproducibility auditor that evaluates repositories for engineering quality, system design patterns, and reproducibility signals using GitHub API analysis.