Yeah, I don't mean stochastically consistent. Semantically consistent. The job of generating content from text and the job of assessing whether two texts represent aligned concepts are two different jobs, and I wouldn't expect a single LLM to do both within itself. That's why you want a second checker.