Sampling strategy · Expert agreement metrics · All parameters auto-determined from corpus · Use CSV for large files (>20 MB)
What would you like to do?
Relationship between LLM sentiment annotation and reviewer star rating. Results on the full corpus (~220k rows) are statistically robust.
Agreement between two independent annotation attempts by the same expert on the same sample — distinct from within-session consistency above (fragment-level self-contradiction) and from Inter-rater Agreement (different people). Only shown for experts with more than one loaded attempt.
Intraclass Correlation Coefficient (two-way random, absolute agreement) on the number of positive/negative/neutral fragments per annotator pair, conditional on both annotators having the dimension active. Thresholds: <0.50 Poor · 0.50–0.74 Moderate · 0.75–0.89 Good · ≥0.90 Excellent (Koo & Mae, 2016).
Convergent validity check: relationship between LLM sentiment annotation and reviewer star rating. A high correlation reflects consistency between the reviewer's numeric rating and the LLM's semantic annotation — not a causal claim.
Sessions are saved on the server and can be resumed at any time. Each expert should use a unique session ID.
Export your completed annotations for use in the metrics computation module.