GLM-5.2 results were sus, so I looked into how the models post-train
and it's slop
the results would be useless in the real world
it's just another benchmark that GLM bros hillclimbed
mind you, GLM-5 was in 22nd place and then a few months later it's suddenly in 1st
part of
This is a new paradigm for interacting with Claude that is significantly more "inline" with all the other human activity org-wide. Once you do all of the under the hood engineering work to make this "just work" (e.g. across tools, integrations, compute environments, memory,