I was using 3.5 Flash Lite (Gemini) , so you are saying that start with a set of questions that represent different classes of the eval set (could be business domain , or query profile like multi hop joins , simple aggreagates) and then let the model generalise and continuously improve your semantic model using real user interactions
Regarding the LLM model , in your experiments have you tried this with non-frontier models because I feel like the more powerful the model the less handholding it needs whereas with a weaker model you need more scaffolding around it to make it reliably good at this stuff