You will evaluate AI conversations for clinical safety and quality while identifying potential risks and model failure modes. Additionally, you will help develop evaluation criteria and scoring rubrics to guide appropriate AI behavior in sensitive mental health contexts.