R
Best cheap way to evaluate LLM outputs without a full eval framework?
Building evals feels like a second product. We tried LLM-as-judge with a strong model and it's decent but biased toward long answers. Anyone using lightweight checks + sampling? What's your minimal viable eval?
댓글 0
첫 댓글을 남겨보세요.
로그인 후 댓글을 쓸 수 있어요.