{
  "title": "What counts as an evaluated result? — Research & questions",
  "text": "MODAVIS (2026). What counts as an evaluated result? — Research & questions. Research notebooks, September 2026 · research guides edition. https://modavis.org/editions/2026-09-r4/research/evaluation/",
  "bibtex": "@misc{modavis_research_evaluation_2026_09_r4,\n  author = {{MODAVIS}},\n  title = {What counts as an evaluated result? — Research & questions},\n  year = {2026},\n  month = {09},\n  version = {2026-09-r4},\n  url = {https://modavis.org/editions/2026-09-r4/research/evaluation/},\n  note = {Content SHA-256: 403e5308e599518077d19cd28b6621820b4b5d800a0db9d4e50dced8f8177d68}\n}",
  "url": "https://modavis.org/editions/2026-09-r4/research/evaluation/",
  "sha256": "403e5308e599518077d19cd28b6621820b4b5d800a0db9d4e50dced8f8177d68",
  "edition": "2026-09-r4",
  "date": "2026-09-26",
  "canonicalPayload": "{\"book\":{\"number\":\"02\",\"slug\":\"research\",\"title\":\"Research & questions\"},\"chapter\":{\"blocks\":[{\"id\":\"six-dimensions\",\"text\":\"Six questions, rather than a single score\",\"type\":\"heading\"},{\"text\":\"Section 8.1.1 describes a quality vector with six dimensions. It is a framework for making judgments explicit, not a calibrated universal ranking. Each dimension needs its own observations, comparison conditions and uncertainty. A missing assessment means “not investigated”; assigning it zero would confuse missing knowledge with a demonstrated failure.\",\"type\":\"p\"},{\"headers\":[\"Dimension\",\"What to examine\",\"Example of evidence\"],\"rows\":[[\"Evidence grounding\",\"Can a representation or assertion be traced to appropriate sources?\",\"Identified source fragment, acquisition context, derivation and review decision.\"],[\"Measurement and model validity\",\"Does the measurement or model address the stated property?\",\"Calibration, uncertainty budget, reference data and an independent comparison.\"],[\"Semantic consistency\",\"Do identities, relations and constraints agree?\",\"Schema and SHACL reports, resolved references and consistent contextual scope.\"],[\"Computational reproducibility\",\"Can the stated transformation be repeated or inspected?\",\"Exact input bytes, tool and parameter versions, outputs and execution records.\"],[\"Experience and interaction\",\"Does the interface support the intended task?\",\"A task-specific study of users, interaction, workload or listening judgments.\"],[\"Access and reuse\",\"Can others obtain and use the relevant material?\",\"Identified distributions, formats, dependencies, rights information and working access routes.\"]],\"type\":\"table\"},{\"text\":\"The dimensions need not be independent, and their numbers are not automatically comparable. A weighted total could hide an essential failure: impressive playback cannot compensate for an untraceable historical claim. Define mandatory conditions first, then compare alternatives on the dimensions appropriate to the question. A classroom demonstration and a metrological comparison legitimately require different evidence.\",\"type\":\"p\"},{\"id\":\"checks-and-claims\",\"text\":\"Match each check to the claim it supports\",\"type\":\"heading\"},{\"text\":\"A SHA-256 match establishes the identity of supplied bytes relative to an expected digest. It does not establish that a microphone was calibrated or a builder attribution is correct. A schema check establishes a specified structural contract. It does not show that a modeled room sounds like the measured room. A reproducible program can reproduce a mistaken assumption perfectly. These checks are valuable because their scope is precise.\",\"type\":\"p\"},{\"text\":\"For an acoustic comparison, specify the reference signal or measurement, alignment, level treatment, bandwidth, sample selection and error statistic. For a timing claim, define the start and end events, device chain, measurement method, trial count and uncertainty. A display refresh rate is not an audio-latency measurement. For a historical claim, inspect source dependence and the instrument state to which each account applies.\",\"type\":\"p\"},{\"id\":\"prospective-studies\",\"text\":\"Studies prepared, but not completed\",\"type\":\"heading\"},{\"text\":\"Section 8.5.2 explicitly identifies five prospective study areas. The dissertation prepares MUSHRA listening comparisons, task-based SUS and NASA-TLX interface studies, external expert panels, apparatus-based action calibration with IMU sensors, and PAMT field validation with uncertainty budgets. These are research plans. They must not be cited as completed listener, usability, expert-panel or field-validation results.\",\"type\":\"p\"},{\"text\":\"The distinction does not erase the technical checks and corpus or signal analyses that were performed. It makes their contribution interpretable. Development-time listening can identify a troublesome artifact and guide a revision; it does not estimate a population’s perceptual preference. A small diagnostic set can reveal a software failure; it does not establish general accuracy over all instruments or collections.\",\"type\":\"p\"},{\"id\":\"independent-comparison\",\"text\":\"Keep optimization separate from confirmation\",\"type\":\"heading\"},{\"text\":\"A reconstruction can become self-confirming if its free parameters are repeatedly adjusted until they match the same observations later used to validate it. The thesis calls for separating evidence-based conditions, optimization objectives and independent checks. Document which observations shaped a model and reserve other suitable observations for evaluation where possible. If no independent comparison exists, describe the result as a model-supported hypothesis with an explicit domain of use.\",\"type\":\"p\"},{\"text\":\"State the claim; identify the object and version; describe the comparison and sample; report results and uncertainty; retain failures; say which checks were not performed. This is a suggested reporting structure derived from the thesis, not an additional VAO conformance profile.\",\"title\":\"A useful evaluation record\",\"type\":\"note\"},{\"items\":[{\"detail\":\"Quantities, denominators and unresolved evidence.\",\"href\":\"/notebooks/research/case-studies/\",\"label\":\"Inspect the reported case-study results\"},{\"detail\":\"The separate layers of a VAO validation claim.\",\"href\":\"/notebooks/standards/conformance/\",\"label\":\"Check technical conformance\"}],\"type\":\"links\"}],\"intro\":\"A research object can be technically valid, scientifically limited and useful for a particular purpose at the same time. The thesis makes those judgments separately.\",\"slug\":\"evaluation\",\"sources\":[{\"detail\":\"Dissertation submitted to Universität Leipzig, 11 September 2026. §8.1.1, pp. 254–256; §8.5, pp. 280–282. Page numbers refer to the printed manuscript pagination. The manuscript is not distributed by this website.\",\"label\":\"Dominik Ukolov · Musikinstrumente im virtuellen Raum (2026)\"},{\"detail\":\"Normative standard, profile index, Dynamic Delivery Profile and conformance specification in the versioned release.\",\"href\":\"https://doi.org/10.5281/zenodo.22214248\",\"label\":\"VAO Standard 0.5.0\"}],\"title\":\"What counts as an evaluated result?\"},\"date\":\"2026-09-26\",\"edition\":\"2026-09-r4\",\"figures\":{}}",
  "payload": {
    "edition": "2026-09-r4",
    "date": "2026-09-26",
    "book": {
      "slug": "research",
      "title": "Research & questions",
      "number": "02"
    },
    "chapter": {
      "slug": "evaluation",
      "title": "What counts as an evaluated result?",
      "intro": "A research object can be technically valid, scientifically limited and useful for a particular purpose at the same time. The thesis makes those judgments separately.",
      "blocks": [
        {
          "type": "heading",
          "id": "six-dimensions",
          "text": "Six questions, rather than a single score"
        },
        {
          "type": "p",
          "text": "Section 8.1.1 describes a quality vector with six dimensions. It is a framework for making judgments explicit, not a calibrated universal ranking. Each dimension needs its own observations, comparison conditions and uncertainty. A missing assessment means “not investigated”; assigning it zero would confuse missing knowledge with a demonstrated failure."
        },
        {
          "type": "table",
          "headers": [
            "Dimension",
            "What to examine",
            "Example of evidence"
          ],
          "rows": [
            [
              "Evidence grounding",
              "Can a representation or assertion be traced to appropriate sources?",
              "Identified source fragment, acquisition context, derivation and review decision."
            ],
            [
              "Measurement and model validity",
              "Does the measurement or model address the stated property?",
              "Calibration, uncertainty budget, reference data and an independent comparison."
            ],
            [
              "Semantic consistency",
              "Do identities, relations and constraints agree?",
              "Schema and SHACL reports, resolved references and consistent contextual scope."
            ],
            [
              "Computational reproducibility",
              "Can the stated transformation be repeated or inspected?",
              "Exact input bytes, tool and parameter versions, outputs and execution records."
            ],
            [
              "Experience and interaction",
              "Does the interface support the intended task?",
              "A task-specific study of users, interaction, workload or listening judgments."
            ],
            [
              "Access and reuse",
              "Can others obtain and use the relevant material?",
              "Identified distributions, formats, dependencies, rights information and working access routes."
            ]
          ]
        },
        {
          "type": "p",
          "text": "The dimensions need not be independent, and their numbers are not automatically comparable. A weighted total could hide an essential failure: impressive playback cannot compensate for an untraceable historical claim. Define mandatory conditions first, then compare alternatives on the dimensions appropriate to the question. A classroom demonstration and a metrological comparison legitimately require different evidence."
        },
        {
          "type": "heading",
          "id": "checks-and-claims",
          "text": "Match each check to the claim it supports"
        },
        {
          "type": "p",
          "text": "A SHA-256 match establishes the identity of supplied bytes relative to an expected digest. It does not establish that a microphone was calibrated or a builder attribution is correct. A schema check establishes a specified structural contract. It does not show that a modeled room sounds like the measured room. A reproducible program can reproduce a mistaken assumption perfectly. These checks are valuable because their scope is precise."
        },
        {
          "type": "p",
          "text": "For an acoustic comparison, specify the reference signal or measurement, alignment, level treatment, bandwidth, sample selection and error statistic. For a timing claim, define the start and end events, device chain, measurement method, trial count and uncertainty. A display refresh rate is not an audio-latency measurement. For a historical claim, inspect source dependence and the instrument state to which each account applies."
        },
        {
          "type": "heading",
          "id": "prospective-studies",
          "text": "Studies prepared, but not completed"
        },
        {
          "type": "p",
          "text": "Section 8.5.2 explicitly identifies five prospective study areas. The dissertation prepares MUSHRA listening comparisons, task-based SUS and NASA-TLX interface studies, external expert panels, apparatus-based action calibration with IMU sensors, and PAMT field validation with uncertainty budgets. These are research plans. They must not be cited as completed listener, usability, expert-panel or field-validation results."
        },
        {
          "type": "p",
          "text": "The distinction does not erase the technical checks and corpus or signal analyses that were performed. It makes their contribution interpretable. Development-time listening can identify a troublesome artifact and guide a revision; it does not estimate a population’s perceptual preference. A small diagnostic set can reveal a software failure; it does not establish general accuracy over all instruments or collections."
        },
        {
          "type": "heading",
          "id": "independent-comparison",
          "text": "Keep optimization separate from confirmation"
        },
        {
          "type": "p",
          "text": "A reconstruction can become self-confirming if its free parameters are repeatedly adjusted until they match the same observations later used to validate it. The thesis calls for separating evidence-based conditions, optimization objectives and independent checks. Document which observations shaped a model and reserve other suitable observations for evaluation where possible. If no independent comparison exists, describe the result as a model-supported hypothesis with an explicit domain of use."
        },
        {
          "type": "note",
          "title": "A useful evaluation record",
          "text": "State the claim; identify the object and version; describe the comparison and sample; report results and uncertainty; retain failures; say which checks were not performed. This is a suggested reporting structure derived from the thesis, not an additional VAO conformance profile."
        },
        {
          "type": "links",
          "items": [
            {
              "label": "Inspect the reported case-study results",
              "href": "/notebooks/research/case-studies/",
              "detail": "Quantities, denominators and unresolved evidence."
            },
            {
              "label": "Check technical conformance",
              "href": "/notebooks/standards/conformance/",
              "detail": "The separate layers of a VAO validation claim."
            }
          ]
        }
      ],
      "sources": [
        {
          "label": "Dominik Ukolov · Musikinstrumente im virtuellen Raum (2026)",
          "detail": "Dissertation submitted to Universität Leipzig, 11 September 2026. §8.1.1, pp. 254–256; §8.5, pp. 280–282. Page numbers refer to the printed manuscript pagination. The manuscript is not distributed by this website."
        },
        {
          "label": "VAO Standard 0.5.0",
          "href": "https://doi.org/10.5281/zenodo.22214248",
          "detail": "Normative standard, profile index, Dynamic Delivery Profile and conformance specification in the versioned release."
        }
      ]
    },
    "figures": {}
  }
}