Howdy! Great questions!
Generalization is definitely still a hard problem. In our experiment, we used a pretty representative sample of the question set. Our goal was actually to have a training question for each "domain" of the set.
When going into production, I think the highest ROI is really to increase that training set so that it covers a large fraction of the types of questions that people ask. Then if someone asks a question that didn't get a good answer, you'd want a feedback mechanism where you get notified when that happens so you can add more training questions.
What model were you using for your evals? It does require a fairly new model for this to work well, but it can be a small new model.