# Dhairyashil R. G. > I find out where the time, the cost and the failures actually go, in compilers, parallel code and AI agents, and I publish the measurement so someone else can re-run it. Senior Member of Technical Staff, AMD (uProf profiler). PhD, IISc Bangalore. Every number on this site has a source, a date and, where one exists, a command to re-run it. Numbers last checked 2026-09-23. ## Who - [About](https://dhairyashilrg.dev/about.md) - [FAQ](https://dhairyashilrg.dev/faq.md) - [CV, typed](https://dhairyashilrg.dev/cv.json) ## Claims worth checking - 14 → 0 — llvm.masked.* calls, PolyBench jacobi-2d (MEDIUM, N=250, TSTEPS=100), vector width 4. https://github.com/llvm/llvm-project/pull/215340 - 317 → 85 — kernel instructions; branches 108 → 12; transfers proved in-bounds 0 → 12. https://github.com/llvm/llvm-project/pull/215340 - 14.70 → 2.89 ms — wall clock, median of 21 runs, −80.3 % (≈5.1×); noise floor 0.45 % (A-vs-A plus the larger MAD); outputs identical. https://github.com/llvm/llvm-project/pull/215340 - 8 pull requests, 5 merged — llvm-project 7 (4 merged), spack-packages 1 (merged); closed-unmerged PRs are not counted. src/data/prs.json (tools/fetch_prs.py) - 64 GFLOP/s — single core, generated matmul + bias + activation kernel. résumé, "ML Compiler and Inference Performance (2026)" - 2.4× — vs a hand-tuned blocked C baseline; 84 % of the measured machine ceiling. résumé, "ML Compiler and Inference Performance (2026)" - ≈6× — same tiled, fused, vectorised kernel; outerproduct vs default vector.contract lowering. résumé, "Codegen analysis of the production MLIR lowering stack" - 2× — DSMC simulator runtime vs Sandia SPARTA, same hardware and problem setup. SankhyaSutra Labs internal benchmark (2018–2024), résumé - 100+ — cluster nodes the simulator scaled to. résumé, SankhyaSutra Labs - 16/36 → 36/36 — DeepSeek V4 Flash, keep-everything vs summaries with pinned rows, question asked first. next_series/T01_context_and_memory/SUMMARY.md in dhairyashilRG/agent-harnesses-2026 - 2 to 35 of 97 — successful injections across five models, undefended, AgentDojo important_instructions. next_series/T02_agent_security/series/FACTS.md in dhairyashilRG/agent-harnesses-2026 - agreement 0.72–0.78, κ 0.45–0.56 — three LLM judges vs AgentDojo ground truth, 300 items. next_series/T03_evals/series/FACTS.md in dhairyashilRG/agent-harnesses-2026 - 1,512 — graded local attempts, eight open models, one MacBook Pro (M3 Max, 128 GB). next_series/T04_local_models/SUMMARY.md in dhairyashilRG/agent-harnesses-2026 - 9 — posts by others on the LLVM Discourse RFC (16 posts, 6 participants, 2026-08-25 to 2026-09-21). https://discourse.llvm.org/t/91649.json ## Work - [Pull requests and results](https://dhairyashilrg.dev/work.md) ## Projects - [Five times faster by proving a load cannot overrun](https://dhairyashilrg.dev/projects/in-bounds.md): Why does a correct, tiled, vectorised kernel spend most of its instructions checking whether each lane is inside the buffer? ## Videos - AI Agent Harnesses (2026) (publishing, 3 of 9 episodes on YouTube): https://github.com/dhairyashilRG/agent-harnesses-2026 - What the Agent Keeps (cut, not yet published): https://github.com/dhairyashilRG/agent-harnesses-2026 - Your Agent Will Be Attacked (cut, not yet published): https://github.com/dhairyashilRG/agent-harnesses-2026 - Evals That Don't Lie (cut, not yet published): https://github.com/dhairyashilRG/agent-harnesses-2026 - Eight Models, One Mac (cut, not yet published): https://github.com/dhairyashilRG/agent-harnesses-2026 ## Log - [Dated entries](https://dhairyashilrg.dev/log.md), [RSS](https://dhairyashilrg.dev/feed.xml) ## How to cite Dhairyashil R. G., dhairyashilrg.dev, , . ## Contact - hello@dhairyashilrg.dev - https://github.com/dhairyashilRG - https://www.linkedin.com/in/dhairyashilrg - https://scholar.google.co.in/citations?user=40V2yM8AAAAJ - https://www.youtube.com/@dhairyashilRG - https://orcid.org/0000-0003-4516-4415 ## Everything - [The whole site as one file](https://dhairyashilrg.dev/llms-full.txt)