{
  "title": "What counts as an evaluated result? — Research & questions",
  "text": "MODAVIS (2026). What counts as an evaluated result? — Research & questions. Research notebooks, September 2026 · Cuntz worked examples edition. https://modavis.org/editions/2026-09-r5/research/evaluation/",
  "bibtex": "@misc{modavis_research_evaluation_2026_09_r5,\n  author = {{MODAVIS}},\n  title = {What counts as an evaluated result? — Research & questions},\n  year = {2026},\n  month = {09},\n  version = {2026-09-r5},\n  url = {https://modavis.org/editions/2026-09-r5/research/evaluation/},\n  note = {Content SHA-256: f8eace2793b229bd05346f2eb62399542f172294eaa9649dc1b6eec669c46807}\n}",
  "url": "https://modavis.org/editions/2026-09-r5/research/evaluation/",
  "sha256": "f8eace2793b229bd05346f2eb62399542f172294eaa9649dc1b6eec669c46807",
  "edition": "2026-09-r5",
  "date": "2026-09-26",
  "canonicalPayload": "{\"book\":{\"number\":\"02\",\"slug\":\"research\",\"title\":\"Research & questions\"},\"chapter\":{\"blocks\":[{\"id\":\"six-dimensions\",\"text\":\"Six questions, rather than a single score\",\"type\":\"heading\"},{\"text\":\"Section 8.1.1 describes a quality vector with six dimensions. It is a framework for making judgments explicit, not a calibrated universal ranking. Each dimension needs its own observations, comparison conditions and uncertainty. A missing assessment means “not investigated”; assigning it zero would confuse missing knowledge with a demonstrated failure.\",\"type\":\"p\"},{\"headers\":[\"Dimension\",\"What to examine\",\"Example of evidence\"],\"rows\":[[\"Evidence grounding\",\"Can a representation or assertion be traced to appropriate sources?\",\"Identified source fragment, acquisition context, derivation and review decision.\"],[\"Measurement and model validity\",\"Does the measurement or model address the stated property?\",\"Calibration, uncertainty budget, reference data and an independent comparison.\"],[\"Semantic consistency\",\"Do identities, relations and constraints agree?\",\"Schema and SHACL reports, resolved references and consistent contextual scope.\"],[\"Computational reproducibility\",\"Can the stated transformation be repeated or inspected?\",\"Exact input bytes, tool and parameter versions, outputs and execution records.\"],[\"Experience and interaction\",\"Does the interface support the intended task?\",\"A task-specific study of users, interaction, workload or listening judgments.\"],[\"Access and reuse\",\"Can others obtain and use the relevant material?\",\"Identified distributions, formats, dependencies, rights information and working access routes.\"]],\"type\":\"table\"},{\"text\":\"The dimensions need not be independent, and their numbers are not automatically comparable. A weighted total could hide an essential failure: impressive playback cannot compensate for an untraceable historical claim. Define mandatory conditions first, then compare alternatives on the dimensions appropriate to the question. A classroom demonstration and a metrological comparison legitimately require different evidence.\",\"type\":\"p\"},{\"id\":\"checks-and-claims\",\"text\":\"Match each check to the claim it supports\",\"type\":\"heading\"},{\"text\":\"A SHA-256 match establishes the identity of supplied bytes relative to an expected digest. It does not establish that a microphone was calibrated or a builder attribution is correct. A schema check establishes a specified structural contract. It does not show that a modeled room sounds like the measured room. A reproducible program can reproduce a mistaken assumption perfectly. These checks are valuable because their scope is precise.\",\"type\":\"p\"},{\"text\":\"For an acoustic comparison, specify the reference signal or measurement, alignment, level treatment, bandwidth, sample selection and error statistic. For a timing claim, define the start and end events, device chain, measurement method, trial count and uncertainty. A display refresh rate is not an audio-latency measurement. For a historical claim, inspect source dependence and the instrument state to which each account applies.\",\"type\":\"p\"},{\"boundary\":\"The notebook’s JSON teaching excerpt is deliberately incomplete and is not presented as a complete, independently conforming VAO package.\",\"concept\":\"Conformance means meeting a specified contract. Different checks address parsing, structure, relationships, required capabilities and file identity.\",\"expanded\":false,\"id\":\"cuntz-conformance\",\"scenario\":\"Take the published observation with value 1281 millimetres. A reader wants to know both whether its record is valid and whether the measurement answers a question about the pipe.\",\"sources\":[{\"detail\":\"Published manifest: identities, measurement observations, realization digests, profiles and rights. Exact examples checked against this release.\",\"href\":\"https://doi.org/10.5281/zenodo.22151203\",\"label\":\"Cuntz Positiv · VAO 0.5.0-rc.2\"},{\"detail\":\"The standard contract is separate from the Cuntz dataset content version 0.5.0-rc.2.\",\"href\":\"https://doi.org/10.5281/zenodo.22214248\",\"label\":\"VAO Standard 0.5.0\"}],\"steps\":[{\"text\":\"A schema can check that the result has the required form and that a numerical field contains a number. Semantic checks can verify that references point to appropriate records.\",\"title\":\"Check the structure\"},{\"text\":\"Resolve the pipe, measurement activity and workbook realization; verify the cited source bytes. These operations establish what the record says and which source it identifies.\",\"title\":\"Check the evidence links\"},{\"text\":\"The workbook timestamp still is not the measurement-session date. A valid record without a documented calibration uncertainty cannot support an uncertainty claim simply because its schema passed.\",\"title\":\"Assess fitness for the question\"}],\"takeaway\":\"Validation and scientific assessment complement one another; each needs an explicitly stated scope.\",\"title\":\"How can a file pass validation but support a limited conclusion?\",\"type\":\"example\"},{\"id\":\"prospective-studies\",\"text\":\"Studies prepared, but not completed\",\"type\":\"heading\"},{\"text\":\"Section 8.5.2 explicitly identifies five prospective study areas. The dissertation prepares MUSHRA listening comparisons, task-based SUS and NASA-TLX interface studies, external expert panels, apparatus-based action calibration with IMU sensors, and PAMT field validation with uncertainty budgets. These are research plans. They must not be cited as completed listener, usability, expert-panel or field-validation results.\",\"type\":\"p\"},{\"text\":\"The distinction does not erase the technical checks and corpus or signal analyses that were performed. It makes their contribution interpretable. Development-time listening can identify a troublesome artifact and guide a revision; it does not estimate a population’s perceptual preference. A small diagnostic set can reveal a software failure; it does not establish general accuracy over all instruments or collections.\",\"type\":\"p\"},{\"id\":\"independent-comparison\",\"text\":\"Keep optimization separate from confirmation\",\"type\":\"heading\"},{\"text\":\"A reconstruction can become self-confirming if its free parameters are repeatedly adjusted until they match the same observations later used to validate it. The thesis calls for separating evidence-based conditions, optimization objectives and independent checks. Document which observations shaped a model and reserve other suitable observations for evaluation where possible. If no independent comparison exists, describe the result as a model-supported hypothesis with an explicit domain of use.\",\"type\":\"p\"},{\"text\":\"State the claim; identify the object and version; describe the comparison and sample; report results and uncertainty; retain failures; say which checks were not performed. This is a suggested reporting structure derived from the thesis, not an additional VAO conformance profile.\",\"title\":\"A useful evaluation record\",\"type\":\"note\"},{\"items\":[{\"detail\":\"Quantities, denominators and unresolved evidence.\",\"href\":\"/notebooks/research/case-studies/\",\"label\":\"Inspect the reported case-study results\"},{\"detail\":\"The separate layers of a VAO validation claim.\",\"href\":\"/notebooks/standards/conformance/\",\"label\":\"Check technical conformance\"}],\"type\":\"links\"},{\"boundary\":\"No new simulation, measured frequency or complete pipe reconstruction is supplied by this example.\",\"concept\":\"A reconstruction combines evidence with assumptions to produce a possible representation. A simulation computes behavior under specified model conditions.\",\"expanded\":false,\"id\":\"cuntz-reconstruction\",\"scenario\":\"Suppose a researcher uses the Cuntz GesL observation to help constrain a pipe model. This is an illustrative reconstruction task; the single value is not a complete geometric description.\",\"sources\":[{\"detail\":\"Published manifest: identities, measurement observations, realization digests, profiles and rights. Exact examples checked against this release.\",\"href\":\"https://doi.org/10.5281/zenodo.22151203\",\"label\":\"Cuntz Positiv · VAO 0.5.0-rc.2\"},{\"detail\":\"Submitted dissertation, 11 September 2026; §§8.2.2.1 and 8.4.1, printed pp. 264–265 and 275–276. The historical description and digital inventory require a resolved mapping; the manuscript is not redistributed here.\",\"label\":\"Dominik Ukolov · Musikinstrumente im virtuellen Raum (2026)\"}],\"steps\":[{\"text\":\"The observation identifies one property, one pipe, one result and one source. It does not provide every interior dimension, material parameter or boundary condition needed by an acoustic model.\",\"title\":\"List what is supplied\"},{\"text\":\"Record the additional dimensions and physical assumptions with their origin. Parameters chosen for convenience should not appear as measured values.\",\"title\":\"Label what is assumed\"},{\"text\":\"If parameters are adjusted to match a known frequency, that same match is not an independent validation. Compare with suitable independent evidence and retain the model’s domain of use.\",\"title\":\"Keep fitting and testing separate\"}],\"takeaway\":\"A plausible reconstruction is a documented hypothesis constrained by evidence, not a recovered certainty.\",\"title\":\"When does a modeled pipe become a historical hypothesis?\",\"type\":\"example\"}],\"intro\":\"A research object can be technically valid, scientifically limited and useful for a particular purpose at the same time. The thesis makes those judgments separately.\",\"slug\":\"evaluation\",\"sources\":[{\"detail\":\"Dissertation submitted to Universität Leipzig, 11 September 2026. §8.1.1, pp. 254–256; §8.5, pp. 280–282. Page numbers refer to the printed manuscript pagination. The manuscript is not distributed by this website.\",\"label\":\"Dominik Ukolov · Musikinstrumente im virtuellen Raum (2026)\"},{\"detail\":\"Normative standard, profile index, Dynamic Delivery Profile and conformance specification in the versioned release.\",\"href\":\"https://doi.org/10.5281/zenodo.22214248\",\"label\":\"VAO Standard 0.5.0\"}],\"title\":\"What counts as an evaluated result?\"},\"date\":\"2026-09-26\",\"edition\":\"2026-09-r5\",\"figures\":{}}",
  "payload": {
    "edition": "2026-09-r5",
    "date": "2026-09-26",
    "book": {
      "slug": "research",
      "title": "Research & questions",
      "number": "02"
    },
    "chapter": {
      "slug": "evaluation",
      "title": "What counts as an evaluated result?",
      "intro": "A research object can be technically valid, scientifically limited and useful for a particular purpose at the same time. The thesis makes those judgments separately.",
      "blocks": [
        {
          "type": "heading",
          "id": "six-dimensions",
          "text": "Six questions, rather than a single score"
        },
        {
          "type": "p",
          "text": "Section 8.1.1 describes a quality vector with six dimensions. It is a framework for making judgments explicit, not a calibrated universal ranking. Each dimension needs its own observations, comparison conditions and uncertainty. A missing assessment means “not investigated”; assigning it zero would confuse missing knowledge with a demonstrated failure."
        },
        {
          "type": "table",
          "headers": [
            "Dimension",
            "What to examine",
            "Example of evidence"
          ],
          "rows": [
            [
              "Evidence grounding",
              "Can a representation or assertion be traced to appropriate sources?",
              "Identified source fragment, acquisition context, derivation and review decision."
            ],
            [
              "Measurement and model validity",
              "Does the measurement or model address the stated property?",
              "Calibration, uncertainty budget, reference data and an independent comparison."
            ],
            [
              "Semantic consistency",
              "Do identities, relations and constraints agree?",
              "Schema and SHACL reports, resolved references and consistent contextual scope."
            ],
            [
              "Computational reproducibility",
              "Can the stated transformation be repeated or inspected?",
              "Exact input bytes, tool and parameter versions, outputs and execution records."
            ],
            [
              "Experience and interaction",
              "Does the interface support the intended task?",
              "A task-specific study of users, interaction, workload or listening judgments."
            ],
            [
              "Access and reuse",
              "Can others obtain and use the relevant material?",
              "Identified distributions, formats, dependencies, rights information and working access routes."
            ]
          ]
        },
        {
          "type": "p",
          "text": "The dimensions need not be independent, and their numbers are not automatically comparable. A weighted total could hide an essential failure: impressive playback cannot compensate for an untraceable historical claim. Define mandatory conditions first, then compare alternatives on the dimensions appropriate to the question. A classroom demonstration and a metrological comparison legitimately require different evidence."
        },
        {
          "type": "heading",
          "id": "checks-and-claims",
          "text": "Match each check to the claim it supports"
        },
        {
          "type": "p",
          "text": "A SHA-256 match establishes the identity of supplied bytes relative to an expected digest. It does not establish that a microphone was calibrated or a builder attribution is correct. A schema check establishes a specified structural contract. It does not show that a modeled room sounds like the measured room. A reproducible program can reproduce a mistaken assumption perfectly. These checks are valuable because their scope is precise."
        },
        {
          "type": "p",
          "text": "For an acoustic comparison, specify the reference signal or measurement, alignment, level treatment, bandwidth, sample selection and error statistic. For a timing claim, define the start and end events, device chain, measurement method, trial count and uncertainty. A display refresh rate is not an audio-latency measurement. For a historical claim, inspect source dependence and the instrument state to which each account applies."
        },
        {
          "type": "example",
          "id": "cuntz-conformance",
          "title": "How can a file pass validation but support a limited conclusion?",
          "concept": "Conformance means meeting a specified contract. Different checks address parsing, structure, relationships, required capabilities and file identity.",
          "scenario": "Take the published observation with value 1281 millimetres. A reader wants to know both whether its record is valid and whether the measurement answers a question about the pipe.",
          "steps": [
            {
              "title": "Check the structure",
              "text": "A schema can check that the result has the required form and that a numerical field contains a number. Semantic checks can verify that references point to appropriate records."
            },
            {
              "title": "Check the evidence links",
              "text": "Resolve the pipe, measurement activity and workbook realization; verify the cited source bytes. These operations establish what the record says and which source it identifies."
            },
            {
              "title": "Assess fitness for the question",
              "text": "The workbook timestamp still is not the measurement-session date. A valid record without a documented calibration uncertainty cannot support an uncertainty claim simply because its schema passed."
            }
          ],
          "takeaway": "Validation and scientific assessment complement one another; each needs an explicitly stated scope.",
          "boundary": "The notebook’s JSON teaching excerpt is deliberately incomplete and is not presented as a complete, independently conforming VAO package.",
          "sources": [
            {
              "label": "Cuntz Positiv · VAO 0.5.0-rc.2",
              "href": "https://doi.org/10.5281/zenodo.22151203",
              "detail": "Published manifest: identities, measurement observations, realization digests, profiles and rights. Exact examples checked against this release."
            },
            {
              "label": "VAO Standard 0.5.0",
              "href": "https://doi.org/10.5281/zenodo.22214248",
              "detail": "The standard contract is separate from the Cuntz dataset content version 0.5.0-rc.2."
            }
          ],
          "expanded": false
        },
        {
          "type": "heading",
          "id": "prospective-studies",
          "text": "Studies prepared, but not completed"
        },
        {
          "type": "p",
          "text": "Section 8.5.2 explicitly identifies five prospective study areas. The dissertation prepares MUSHRA listening comparisons, task-based SUS and NASA-TLX interface studies, external expert panels, apparatus-based action calibration with IMU sensors, and PAMT field validation with uncertainty budgets. These are research plans. They must not be cited as completed listener, usability, expert-panel or field-validation results."
        },
        {
          "type": "p",
          "text": "The distinction does not erase the technical checks and corpus or signal analyses that were performed. It makes their contribution interpretable. Development-time listening can identify a troublesome artifact and guide a revision; it does not estimate a population’s perceptual preference. A small diagnostic set can reveal a software failure; it does not establish general accuracy over all instruments or collections."
        },
        {
          "type": "heading",
          "id": "independent-comparison",
          "text": "Keep optimization separate from confirmation"
        },
        {
          "type": "p",
          "text": "A reconstruction can become self-confirming if its free parameters are repeatedly adjusted until they match the same observations later used to validate it. The thesis calls for separating evidence-based conditions, optimization objectives and independent checks. Document which observations shaped a model and reserve other suitable observations for evaluation where possible. If no independent comparison exists, describe the result as a model-supported hypothesis with an explicit domain of use."
        },
        {
          "type": "note",
          "title": "A useful evaluation record",
          "text": "State the claim; identify the object and version; describe the comparison and sample; report results and uncertainty; retain failures; say which checks were not performed. This is a suggested reporting structure derived from the thesis, not an additional VAO conformance profile."
        },
        {
          "type": "links",
          "items": [
            {
              "label": "Inspect the reported case-study results",
              "href": "/notebooks/research/case-studies/",
              "detail": "Quantities, denominators and unresolved evidence."
            },
            {
              "label": "Check technical conformance",
              "href": "/notebooks/standards/conformance/",
              "detail": "The separate layers of a VAO validation claim."
            }
          ]
        },
        {
          "type": "example",
          "id": "cuntz-reconstruction",
          "title": "When does a modeled pipe become a historical hypothesis?",
          "concept": "A reconstruction combines evidence with assumptions to produce a possible representation. A simulation computes behavior under specified model conditions.",
          "scenario": "Suppose a researcher uses the Cuntz GesL observation to help constrain a pipe model. This is an illustrative reconstruction task; the single value is not a complete geometric description.",
          "steps": [
            {
              "title": "List what is supplied",
              "text": "The observation identifies one property, one pipe, one result and one source. It does not provide every interior dimension, material parameter or boundary condition needed by an acoustic model."
            },
            {
              "title": "Label what is assumed",
              "text": "Record the additional dimensions and physical assumptions with their origin. Parameters chosen for convenience should not appear as measured values."
            },
            {
              "title": "Keep fitting and testing separate",
              "text": "If parameters are adjusted to match a known frequency, that same match is not an independent validation. Compare with suitable independent evidence and retain the model’s domain of use."
            }
          ],
          "takeaway": "A plausible reconstruction is a documented hypothesis constrained by evidence, not a recovered certainty.",
          "boundary": "No new simulation, measured frequency or complete pipe reconstruction is supplied by this example.",
          "sources": [
            {
              "label": "Cuntz Positiv · VAO 0.5.0-rc.2",
              "href": "https://doi.org/10.5281/zenodo.22151203",
              "detail": "Published manifest: identities, measurement observations, realization digests, profiles and rights. Exact examples checked against this release."
            },
            {
              "label": "Dominik Ukolov · Musikinstrumente im virtuellen Raum (2026)",
              "detail": "Submitted dissertation, 11 September 2026; §§8.2.2.1 and 8.4.1, printed pp. 264–265 and 275–276. The historical description and digital inventory require a resolved mapping; the manuscript is not redistributed here."
            }
          ],
          "expanded": false
        }
      ],
      "sources": [
        {
          "label": "Dominik Ukolov · Musikinstrumente im virtuellen Raum (2026)",
          "detail": "Dissertation submitted to Universität Leipzig, 11 September 2026. §8.1.1, pp. 254–256; §8.5, pp. 280–282. Page numbers refer to the printed manuscript pagination. The manuscript is not distributed by this website."
        },
        {
          "label": "VAO Standard 0.5.0",
          "href": "https://doi.org/10.5281/zenodo.22214248",
          "detail": "Normative standard, profile index, Dynamic Delivery Profile and conformance specification in the versioned release."
        }
      ]
    },
    "figures": {}
  }
}