Use reasoning during training while omitting it during execution to preserve the resulting policy improvements without paying the reasoning latency cost.
Soon you can unlock how this was done.
Behind this:
the tool used · the method · what actually resulted · the manual work it replaced.