Inference AIops

Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.

Otherv0.6.0