Analyst workspace

Create an account or sign in to evaluate the engines. Visitors can browse all public evaluations and current accuracy results without an account.

Not signed inMy reviews

Guidelines · Purpose and use

This webpage allows analysts to evaluate 335 compositions from the When in Rome corpus and rate Roman-numeral results produced by musWM, AnalysisGNN, AugmentedNet, and Dusk Audio. Analysts create an account before scoring; each completed evaluation is publicly visible, and the accumulated ratings determine each engine's overall accuracy.

  1. Select a work from Benchmark workbook. Its analysis workbook and matching MusicXML score will be loaded together when you press Open workbook and score.
  2. Review how all four engines analyzed each specific harmonic event across the work's measures and beats. The corresponding notated measure is displayed directly above the analysis rows.
  3. After examining a measure, score each Roman-numeral prediction according to the Validation principles. These ratings contribute to the aggregate benchmark; your individual review remains visible in your account and to the administrator.
  4. Use Previous and Next to move between measures, then complete and download your reviewed workbook when finished.

Benchmark Measure Reviewer

Compare event-level Roman-numeral predictions from musWM, AnalysisGNN, AugmentedNet, and Dusk Audio against the notes in each corresponding MusicXML measure. Reviewers score harmonic function, inversion, and triad/tetrad agreement, with measure-balanced accuracy calculated for every engine.

musWM weighted accuracy0 scored
AnalysisGNN weighted accuracy0 scored
AugmentedNET weighted accuracy0 scored
Dusk Audio weighted accuracy*0 scored
Row weight = 5 ÷ number of rows in its measure. Empty score cells are excluded.
* Dusk Audio uses musWM global and local key context; key finding is not evaluated.
Load both files to begin.

MusicXML measure

No MusicXML loaded.

Workbook rows

No workbook loaded.
Set selected score