The AutoSOTA paper is highly interesting, Li, et al decided to dispense with having a monolithic model and do a function-based approach, like how CrewAI works. I'd be interested in a counterfactual study: remove many of the guardrails and, with an intelligent enough model, provide it contours into the harness instead of hard stops. Would it revolutionize science? It may. MoE, as used here, is one approach, but I believe the team there over-constrained it because by default MoE is already constrained to a role, My thought experiment: use deterministic force functions to validate the result before accumulating it, acting as the ultimate guardrail, but leverage the non-deterministic LLM's capability to explore. Either way, it's a great paper, highly recommend reading it.